Hugging Face Daily PapersJie Yang, Zhengyu Fang, Zelin Xu2 min readpaperadvanced
LastOPD: Taming Collapse in Latent On-Policy Distillation
Summary
The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.
- Latent supervision boosts early MATH‑500 scores but later collapses, with alignment metrics misleadingly improving.
- Collapse arises because deeper student layers are forced into teacher states they cannot interpret.
- LastOPD applies latent alignment only to the last‑layer representation and crossfades to token‑level OPD within 10 steps, preventing collapse.
- Experiments show +5.55 / +4.02 points over token‑only OPD for 4B and 8B teachers and reach comparable final performance in ~50% of the steps.
LLM engineers using distillation need a stable, efficient recipe; LastOPD offers a proven way to avoid latent‑signal collapse and speed up student training.
8/10