SyncAI.news, a Varaisys broadcasting
Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs
YW

Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng, Xiang Chen

· 1 min read

ResearcharXiv cs.AI

Before the Token Commits: Trajectory-Level Benchmarking of Visual Hallucinations in Diffusion VLMs

arXiv:2609.34772v1 Announce Type: new Abstract: Multimodal diffusion language models generate responses by iteratively unmasking tokens, making each answer the endpoint of a multi-step trajectory rather than an immediate commitment. Hallucination benchmarks built for autoregressive models evaluate only the final output, and therefore cannot determine whether an unsupported claim in diffusion VLMs appears late or has already stabilized before any answer token is revealed. We introduce DynaHall, a trajectory-level benchmark of annotation-backed binary visual propositions covering object existence, counting, attributes, and relations, with controlled hard negatives graded by visual prior. DynaHall is paired with a commitment-aware protocol that records the intermediate answer tendency at every unmasking step alongside the committed output. Across five diffusion VLMs from three architecture families, visual hallucination is settled before commitment: an unsupported answer is already the preferred state while the answer position is still masked, and later unmasking steps rarely reverse it, so the failure is not introduced at the write step. This holds across decoding schedules, answer formats, and open-ended generation. DynaHall also exposes failures hidden by final-output metrics, including counting and relation collapse, prior-driven false positives, and attribute errors whose direction changes by type. Guided by this diagnosis, PGS (Pre-commitment Gradient Steering) edits still-masked answer states to reduce false positives, bringing the affirmation rate close to balance, and transfers to another architecture without degrading general ability. DynaHall and PGS suggest that hallucination should be measured and mitigated along the generation trajectory of diffusion VLMs, not only at the final answer.

Original source

This story was published by arXiv cs.AI and written by Yadong Wang, Siping Yue, Yu Tian, Chuanxing Geng, Xiang Chen. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News