SyncAI.news, a Varaisys broadcasting
Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models
ZW

Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu

· 1 min read

ResearcharXiv cs.AI

Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

arXiv:2609.38025v1 Announce Type: cross Abstract: On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by $9.7$ points over vanilla OPD, and enables the smaller student to surpass its larger teacher.

Original source

This story was published by arXiv cs.AI and written by Zhenyu Wang, Tianze Wang, Linjun Zhang, Yifan Hu. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News