proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersWenze Lin, Jiyuan Long, Jiale Zhao2 min readpaperadvanced

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Summary

On‑policy distillation of LLMs traditionally uses KL divergence, but this paper shows that merely aligning the update direction toward the teacher—via a simple (+1/‑1) token‑wise reward—achieves the same effect. Building on this, they introduce Consensus Multi‑Teacher OPD, which lets every sample learn from all teachers and consistently beats the prior single‑teacher approach on math and code ben…

  • KL divergence can be replaced by a sign‑based reward (+1/-1) that only enforces update direction toward the teacher.
  • Effective OPD only needs correct direction on a small subset of tokens with high teacher‑student disagreement.
  • Consensus Multi‑Teacher OPD (C‑MOPD) supervises each sample with all teachers, avoiding domain conflicts.
  • C‑MOPD consistently outperforms the original single‑teacher MOPD on both math and code benchmarks.

LLM engineers and researchers can simplify loss design and improve multi‑teacher distillation pipelines, gaining better performance with less computational overhead.

8/10

Related reading

  1. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

    On-policy distillation (OPD) can lead to excessively long student responses, a phenomenon called length inflation. This paper identifies "termination-token mismatch" between base students and post-trained teachers as a key source, where models place stopping probability on different EOS tokens. Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigate…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. LastOPD: Taming Collapse in Latent On-Policy Distillation

    The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

    Hugging Face Daily Papersarxiv.org2 minpaper