SyncAI.news, a Varaisys broadcasting
FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models
DT

Darshan Thaker, Lachlan Ewen MacDonald, Ren\'e Vidal

· 1 min read

ResearcharXiv cs.CL

FastGuide: Accelerating Reward Guidance for Diffusion Large Language Models

arXiv:2609.36202v1 Announce Type: new Abstract: Gradient-based reward guidance provides a flexible way to use downstream reward models to control masked diffusion language models at inference time. However, its computational cost remains high as each decoding iteration incurs expensive diffusion model forward passes and reward model backpropagation steps. To address this, we introduce FastGuide, an adaptive hybrid of parallel and autoregressive decoding to accelerate reward guidance for diffusion language models. In analogy to parallel decoding, FastGuide amortizes the cost of reward model backpropagation by computing guidance once per decoding step and reusing it to generate multiple tokens. Within each decoding step, FastGuide makes diffusion forward passes autoregressive by unmasking tokens one at a time while efficiently recomputing token distributions after each unmasking by utilizing KV caching techniques and sparse recomputation of attention. Lastly, to adapt hybrid decoding to the model's confidence, FastGuide defers any token that the model is unconfident about under its recomputed distribution. Experiments on three reward benchmarks demonstrate that FastGuide is up to $4.4\times$ faster than sequential reward-guided decoding while retaining similar generation quality.

Original source

This story was published by arXiv cs.CL and written by Darshan Thaker, Lachlan Ewen MacDonald, Ren\'e Vidal. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News