SyncAI.news, a Varaisys broadcasting
Mentored Decoding: Faster Inference meets Boosting
VT

Vivien Tran-Thien, Richard Nock

· 1 min read

ResearcharXiv cs.LG

Mentored Decoding: Faster Inference meets Boosting

arXiv:2609.30474v1 Announce Type: new Abstract: Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among which (i) the particularly appealing geometric nature of the total variation case, (ii) simple approximations for any $f$-divergence in direct relation with boosting compliance, and (iii) a $\textit{divergence independent}$ $O(n)$ space and $O(\mathrm{sort}(n))$ time data structure built on drafter and target outputs, which allows to query the optimal parameters of the dual problem in $O(\log n)$ time and constructing optimal mentored distributions in $O(n)$ time for any $f$-divergence.

Original source

This story was published by arXiv cs.LG and written by Vivien Tran-Thien, Richard Nock. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News