
XQ
Xinwei Qiang, Xiang Fang, Chang Chen, Zaifeng Pan, Yue Guan, Yufei Ding
· 1 min read
ResearcharXiv cs.CL
Beyond Parallel Blindness: Information Floors and Model Gaps in Block Drafting
arXiv:2608.27339v3 Announce Type: replace-cross
Abstract: Block drafters propose several tokens in one forward pass, before earlier target tokens are realised. Their rejection mixes two losses: missing within-block path information and imperfect modelling of observable information. Accepted length cannot distinguish them. We separate the two with an information floor, the minimum expected rejection at a specified conditioning order; rejection above this floor is the model gap. Estimating both from target rollouts across four domains, four open-weight targets, and a frontier API target yields three findings. First, the all-parallel floor reaches $0.286$ at the final slot on Qwen3-4B, limiting even the best proposal to $71\%$ per-slot acceptance. Second, one realised token removes $86$--$100\%$ of this floor, a locality also recovered by an independent mutual-information analysis. Third, current drafters remain far above their floors: the final-slot model gap accounts for $43$--$64\%$ of DFlash rejection and $85$--$92\%$ of DSpark's oracle-conditioned rejection. These findings separate the value of short-range conditioning from proposal quality. Guided by this distinction, we replace a context-independent predecessor correction with a prefix-attention head that reads committed context conditional on the predecessor. With supporting backbone components, the resulting Qwen3-4B drafter improves mean serving accepted length by $3.83\%$ over the released DSpark checkpoint across nine tasks, without increasing the conditioning order.
Original source
This story was published by arXiv cs.CL and written by Xinwei Qiang, Xiang Fang, Chang Chen, Zaifeng Pan, Yue Guan, Yufei Ding. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


