proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJie Yang, Zhengyu Fang, Zelin Xu2 min readpaperadvanced

LastOPD: Taming Collapse in Latent On-Policy Distillation

Summary

The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

  • Latent supervision boosts early MATH‑500 scores but later collapses, with alignment metrics misleadingly improving.
  • Collapse arises because deeper student layers are forced into teacher states they cannot interpret.
  • LastOPD applies latent alignment only to the last‑layer representation and crossfades to token‑level OPD within 10 steps, preventing collapse.
  • Experiments show +5.55 / +4.02 points over token‑only OPD for 4B and 8B teachers and reach comparable final performance in ~50% of the steps.

LLM engineers using distillation need a stable, efficient recipe; LastOPD offers a proven way to avoid latent‑signal collapse and speed up student training.

8/10

Related reading

  1. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

    On-policy distillation (OPD) can lead to excessively long student responses, a phenomenon called length inflation. This paper identifies "termination-token mismatch" between base students and post-trained teachers as a key source, where models place stopping probability on different EOS tokens. Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigate…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

    The paper presents PMOPD, a projection-based method that tracks low-dimensional subspaces of task-specific parameter updates during multi-teacher on-policy distillation and removes interfering components. Experiments on Qwen2.5-7B and Llama-3.1-8B show consistent 2-point gains across code, reasoning, and math tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

    The paper studies gradient‑estimation noise in sparse on‑policy distillation (OPD) and introduces the Information‑Efficiency Ratio (IER), a signal‑to‑noise based metric for selecting which tokens to supervise. IER is derived from an information‑geometry analysis with an optimal scalar baseline, and a candidate‑set approximation lets it be combined with existing usefulness scores while keeping the…

    Hugging Face Daily Papersarxiv.org1 minpaper