SyncAI.news, a Varaisys broadcasting
Divergence controls entropy in distillation
NZ

Nicolas Zucchet, Scott W. Linderman

· 1 min read

ResearcharXiv cs.CL

Divergence controls entropy in distillation

arXiv:2610.03529v1 Announce Type: cross Abstract: Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, studying how the entropy of the student depends on the data and the divergence that define the distillation objective. We prove that forward KL inflates the entropy of the student above that of the teacher. Since cross-entropy training is a special case, this yields an identity that we verify quantitatively in pretraining and supervised finetuning. Other divergences come with no such guarantee: reverse KL deflates entropy until the gap between student and teacher gets too large, and interpolating between the two changes entropy smoothly early in training but abruptly at convergence. The lower entropy of on-policy distillation comes from token-level reverse KL, not from on-policy sampling. The divergence therefore acts as an implicit entropy regularizer, whose role is clearest in self-distillation: as conditioning on privileged information deflates entropy, the divergence hyperparameters that work best are those that compensate for it.

Original source

This story was published by arXiv cs.CL and written by Nicolas Zucchet, Scott W. Linderman. SyncAI.news shows a preview; the complete article is on the publisher's site.

Read the full story on arxiv.org

Similar News