SyncAI.news, a Varaisys broadcasting
Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions
SL

Siting Liang, Luca Rippe, Omar Adjali, Daniel Sonntag

· 1 min read

ResearcharXiv cs.CL

Evaluating Explanation-Driven Vision-Language Reasoning via Generation Order Interventions

arXiv:2609.29496v1 Announce Type: new Abstract: Natural language explanation generation serves as a key mechanism for exposing and evaluating vision-language reasoning. Prior work on explanation-driven vision-language models predominantly follows a post-hoc (answer-first) paradigm, implicitly suggesting that supervised rationales can reflect underlying reasoning processes. In contrast, modern large vision-language models increasingly exhibit a rationale-first generation tendency, which more closely aligns with structured, stepwise reasoning. In this work, we systematically evaluate whether explanations are causally tied to model predictions within a single generation step under a controlled experimental setup, explicitly eliminating unnecessary chain-of-thought or other intermediate reasoning processes across knowledge-intensive QA, visual entailment, and compositional grounding benchmarks. We find that larger models emerge as a prerequisite for reliably supporting rationale-first reasoning at scale. However, answer-first generation is less prone to format-related errors in structured output. Overall, explanation ordering, model scale and pre-training knowledge, task-specific fine-tuning, and task structure jointly influence both prediction accuracy and reasoning faithfulness.

Original source

This story was published by arXiv cs.CL and written by Siting Liang, Luca Rippe, Omar Adjali, Daniel Sonntag. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News