
KK
Keuntae Kim, Yong Suk Choi
· 1 min read
ResearcharXiv cs.CV
Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
arXiv:2610.00953v1 Announce Type: new
Abstract: An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual question-answering benchmarks, we find that 89.4-98.1% of the rationale-side canvas remains unwritten at stabilization in single-block, EOS-suppressed LaViDa runs. On V*Bench, reducing block length from 128 to 8 changes this fraction from 89.4% to 1.7%, together with answer coverage and the eligible observation window. Under EOS-enabled prompting, direct instructions improve Nemotron's overall accuracy by 15.0 and 19.5 percentage points on M3CoT and ScienceQA, but reduce LaViDa/V*Bench accuracy by 11.0 points. A symmetric decomposition associates the larger absolute component of each change with coverage rather than conditional accuracy. Matched-canvas image ablations measure visual sensitivity alongside answer stabilization, separating the two temporal readouts. Together, these measurements distinguish answer stabilization, rationale unfolding, and visual sensitivity, and identify coverage as the larger component of the prompting differences.
Original source
This story was published by arXiv cs.CV and written by Keuntae Kim, Yong Suk Choi. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


