SyncAI.news, a Varaisys broadcasting
Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation
YS

Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao, Jing Lyu, Dong Liu

· 1 min read

ResearcharXiv cs.AI

Interpolated Policy Distillation: A Controllable Continuum Between Off-Policy and On-Policy Distillation

arXiv:2609.37170v1 Announce Type: cross Abstract: Off-policy and on-policy distillation have traditionally been formulated as separate paradigms, each favoring a different property of distillation trajectories. Teacher-generated (off-policy) traces are typically high-quality but lie far from the student's distribution, whereas student-generated (on-policy) rollouts are more learnable but often contain erroneous reasoning. We view these paradigms as the endpoints of a policy continuum and posit that a more effective rollout policy may lie in between. We introduce \textbf{Interpolated Policy Distillation (IPD)}, which defines the next-token distribution at every decoding step as an explicit linear interpolation between the student and teacher distributions. The interpolation operates at the distribution level, token by token, and its coefficient provides direct control over the balance between trajectory quality and student learnability. Naively sampling from this policy would require sequentially querying the teacher at every token and is thus expensive. To make IPD practical, we accelerate it with a new speculative-decoding rule while exactly preserving the interpolated next-token distribution.At the trajectory level, the resulting rollouts naturally interleave student- and teacher-generated segments. Unlike recent heuristic segment-interleaving methods, however, this interleaving is induced by an exactly realized token-level interpolated policy rather than by hand-designed switching rules. Across text-only and multimodal reasoning benchmarks, IPD consistently outperforms both endpoint policies (SFT and OPD), their conventional two-stage combination (SFT-then-OPD), and recent heuristic segment-interleaving methods, demonstrating that token-level policy interpolation better balances trajectory quality and student learnability.

Original source

This story was published by arXiv cs.AI and written by Youxu Shi, Yifan Sun, Dacheng Yin, Haomiao Tang, Guangting Wang, Fengyun Rao, Jing Lyu, Dong Liu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News