
YS
Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro
· 1 min read
ResearcharXiv cs.CL
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
arXiv:2606.10829v5 Announce Type: replace
Abstract: Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-\(k\), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule that leaves the base sampler's stopping rule unchanged and greedily discounts each token-wise confidence score according to its attention to already selected positions, weighted by their prediction uncertainty. Across LLaDA-8B-Base and Dream-7B-Base on the reasoning benchmarks GSM8K and MATH500 and the code benchmarks HumanEval and MBPP, plugging ADAS into all three samplers improves low-NFE performance at matched denoiser evaluations by \(9.11\) and \(10.46\) percentage points on average, respectively, with \(3.1\%\) per-forward runtime overhead. Code is available at https://github.com/yusufsahin99/ADAS.
Original source
This story was published by arXiv cs.CL and written by Yusuf Sahin, Ahmed Rockey Saikia, Volkan Cevher, Paolo Favaro. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


