
ZN
Zixiang Ni, Zhuo Hu, Renjie Cao, Weijie Ren, Binqin Shi, Weijia Zhang, Shuheng Cao, Zhicheng Shi, Zhenhao Zhang, Haomin Wen, Zhiyuan Hu
· 1 min read
ResearcharXiv cs.LG
ReTaCo: Residual-Target Control for On-Policy Distillation
arXiv:2609.39275v1 Announce Type: new
Abstract: On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-$k$ mass toward one even after the student matches the teacher's relative probabilities within the top-$k$ set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-$k$ tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-$k$ mass $m$, the residual target is $(1-\beta)(1-m)$ for $\beta\in[0,1]$: $\beta=0$ preserves the teacher's mass, and larger $\beta$ moves more mass onto the top-$k$ tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-$k$ mass lies between $m$ and $m+\beta(1-m)$ and increases monotonically with $\beta$; at $\beta=0$, underestimated top-$k$ tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.
Original source
This story was published by arXiv cs.LG and written by Zixiang Ni, Zhuo Hu, Renjie Cao, Weijie Ren, Binqin Shi, Weijia Zhang, Shuheng Cao, Zhicheng Shi, Zhenhao Zhang, Haomin Wen, Zhiyuan Hu. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


