Hugging Face Daily PapersWenze Lin, Jiyuan Long, Jiale Zhao2 min readpaperadvanced
Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
Summary
On‑policy distillation of LLMs traditionally uses KL divergence, but this paper shows that merely aligning the update direction toward the teacher—via a simple (+1/‑1) token‑wise reward—achieves the same effect. Building on this, they introduce Consensus Multi‑Teacher OPD, which lets every sample learn from all teachers and consistently beats the prior single‑teacher approach on math and code ben…
- KL divergence can be replaced by a sign‑based reward (+1/-1) that only enforces update direction toward the teacher.
- Effective OPD only needs correct direction on a small subset of tokens with high teacher‑student disagreement.
- Consensus Multi‑Teacher OPD (C‑MOPD) supervises each sample with all teachers, avoiding domain conflicts.
- C‑MOPD consistently outperforms the original single‑teacher MOPD on both math and code benchmarks.
LLM engineers and researchers can simplify loss design and improve multi‑teacher distillation pipelines, gaining better performance with less computational overhead.
8/10