SyncAI.news, a Varaisys broadcasting
What Keeps Vision-Language Models Looking at the Image?
HO

Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo

· 1 min read

ResearcharXiv cs.AI

What Keeps Vision-Language Models Looking at the Image?

arXiv:2607.12815v3 Announce Type: replace Abstract: When do vision-language models need direct access to the image while generating an answer? We study image dependence during answer generation by examining how the visual information needed for the current question becomes available in context. We intervene on direct image access while retaining previously computed states. Across real-image and synthetic tasks, we show that, depending on the generation process, direct access can continue to support accuracy after question processing. Supplying the required attributes as text in the context weakens this dependence. On synthetic tasks, we also examine how dependence changes as the model itself states the required attributes. Before attribute expression, severing access reduces accuracy, and replacing image-side states shifts answers toward the counterfactual content. After sufficient expression, both interventions have smaller effects. Even with an identical generated prefix, dependence differs according to whether image access was available during question processing. Thus, both the visible text and the preceding image access matter. Several of these patterns hold across model families, including Qwen2.5-VL-32B and InternVL3-14B. These findings offer a view of image dependence in terms of the information needed for the current question and the history of image access, beyond generation position alone. This perspective provides a basis for deciding when to reduce visual access during an answer and which visual information to retain for subsequent questions.

Original source

This story was published by arXiv cs.AI and written by Hiroto Osaka, Shohei Taniguchi, Gouki Minegishi, Kai Yamashita, Masahiro Suzuki, Yutaka Matsuo. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News