SyncAI.news, a Varaisys broadcasting
Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning
PP

Pengcheng Pan, Xinfang Zhang

· 1 min read

ResearcharXiv cs.CL

Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning

arXiv:2608.04452v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) can miss fine details in a full image that they recognize in a closer view. Recovering this evidence requires deciding where to look and how much surrounding context to retain. We present Q-CueGraph, a query-conditioned evidence acquisition method for frozen MLLMs. For text-rich images, it builds a reusable graph of OCR lines and layout relations. Each question activates anchors, expands them into contextual regions, and selects candidates for a single observation window. Query-conditioned object detections support natural-image search through the same region-selection and composition interface. A lightweight candidate scorer further learns which observations support correct answers from frozen-reader feedback and training answers, without evidence-box supervision. Across six benchmarks, we examine the roles of query conditioning, evidence composition, and learned answerability. With Qwen2.5-VL-7B, Q-CueGraph raises V*Bench accuracy from 0.696 to 0.832 using 19.1% of source-image area, and retains 92% of full-image ANLS on InfographicVQA using about half the image area. The analyses show that useful evidence depends on both its relevance to the question and the context available to the reader. Q-CueGraph makes these choices explicit before answer generation.

Original source

This story was published by arXiv cs.CL and written by Pengcheng Pan, Xinfang Zhang. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News