
JD
Jiacheng Du, Weiwei Xie, Tianyi Du, Shaoxiong Guo, Qibing Ren, Jiaheng Zhang
· 1 min read
ResearcharXiv cs.AI
Learning from a Thoughtful Teacher: Adaptive On-Policy Self-Distillation for Mathematical Reasoning
arXiv:2609.32667v1 Announce Type: new
Abstract: On-policy self-distillation (OPSD) trains a question-only student with token-level feedback from a teacher given training-only privileged information (PI). OPSD therefore provides dense, on-policy supervision, and is free of a larger external teacher, but its effectiveness rests on how PI is designed and utilized. Our preliminary diagnostics suggest a significant gap between teacher utility and student learnability, where a small fraction of high-disagreement tokens dominate the distillation signal, and short teacher continuations at these positions further expose more explicit PI leakage than transferable correction cues, indicating a strong intent on injecting PI-conditioned shortcuts. We propose Adaptive On-Policy Self-Distillation (AOPSD), which adapts what information the teacher receives and how strongly its feedback influences learning. AOPSD encodes each solution as a reasoning DAG, orders problems by the student's evolving capability, and reveals only the affordable subgraph and its next frontier as PI. For high-disagreement tokens, AOPSD utilizes short teacher continuations as probes to encourage useful guidance while mitigating PI-conditioned shortcuts among teacher supervisions. On HMMT25, AIME24, AIME25, and BRUMo25, AOPSD achieves 72.5% Pass@8, which is 6.7 percentage points above OPSD and 4.2 above the strongest competing baseline while reducing 15 percentage points of training time at lower cost.
Original source
This story was published by arXiv cs.AI and written by Jiacheng Du, Weiwei Xie, Tianyi Du, Shaoxiong Guo, Qibing Ren, Jiaheng Zhang. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


