SyncAI.news, a Varaisys broadcasting
From Priors to Perception: Grounding Video-LLMs in Physical Reality
ZZ

Zicheng Zhao, Chaofan Gan, Shijie Li, Weiyao Lin

· 1 min read

ResearcharXiv cs.CV

From Priors to Perception: Grounding Video-LLMs in Physical Reality

arXiv:2605.04515v2 Announce Type: replace Abstract: Video Large Language Models (Video-LLMs) excel in general video understanding but often base physical judgments on event expectations rather than observations. We find that they not only rationalize physically impossible events, but also incorrectly report expected outcomes in physically plausible yet counter-intuitive scenarios, despite clear contradictory visual evidence. We provide the first unified account of these failures as Semantic Prior Dominance (SPD): semantic expectations override conflicting visual evidence. To diagnose and overcome these failures, we introduce PriorPair, a high-fidelity paired video benchmark covering both conflicts across eight physical categories. Its category-directed construction pairs prior-aligned events with semantically matched but prior-conflicting counterparts. We further propose Physics-Anchored Reasoner (PhyAR), a lightweight training framework. Its Visually Anchored Reasoning Chain grounds physical attribution and judgment in explicit visual observations, while Paired Data Binding strengthens learning of decisive event differences through matched-pair supervision. Standard LoRA fine-tuning with PhyAR substantially improves physical reasoning and reduces prior-driven judgment bias across different Video-LLM backbones without architectural modifications. These gains extend to an external physical benchmark while largely preserving general video understanding.

Original source

This story was published by arXiv cs.CV and written by Zicheng Zhao, Chaofan Gan, Shijie Li, Weiyao Lin. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News