
ZL
Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo
· 1 min read
ResearcharXiv cs.CL
Recovering General Capabilities via Uncertainty-Calibrated Multi-Teacher On-Policy Distillation
arXiv:2608.26735v2 Announce Type: replace
Abstract: Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities. We study this trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized model learns from domain and general teachers on its own sampled trajectories. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, and advantage sign alone does not establish whether the proposed update direction is reliable. We propose Uncertainty-Calibrated MOPD (UCMOPD), which addresses these limitations through two complementary mechanisms. Golden-Gain Enhancement combines higher-temperature exploration with a standard-temperature anchor and retains trajectories whose positive learning signal matches or exceeds the prompt-specific anchor. Teacher-Endorsement Filtering then uses centered log-likelihood (CLL) to estimate each retained token's plausibility relative to the teacher's uncertainty and probabilistically preserves updates whose directions are supported by that endorsement. Across role-playing and medical-domain specialization, UCMOPD improves the general-capability average over standard MOPD by $4.48\%$ and $7.86\%$, respectively, while maintaining vertical-domain performance. Component ablations and diagnostic analyses support the intended roles of the two mechanisms: exposing and selecting stronger positive signals at the trajectory level and validating update directions through teacher endorsement at the token level.
Original source
This story was published by arXiv cs.CL and written by Ziyuan Liu, Jiao Ou, Jian Liang, Ruiming Tang, Cheng Luo. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


