proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersLanglin Huang, Hao Liu, Mononito Goswami1 min readpaperadvanced

On the Off-Policy Teacher in On-Policy Distillation

Summary

On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

  • OPD teachers degrade when supervising student-generated, off-policy trajectories.
  • Teacher continuation performance worsens as student-generated prefixes lengthen.
  • SCOUT is a co-training framework that adapts the teacher to student prefixes.
  • SCOUT uses RL with verifiable rewards to optimize the teacher's conditional ability.

Engineers working on improving large language model efficiency and performance through distillation will find this relevant for addressing a core limitation in current OPD methods.

8/10

Related reading

  1. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

    On-policy distillation (OPD) can lead to excessively long student responses, a phenomenon called length inflation. This paper identifies "termination-token mismatch" between base students and post-trained teachers as a key source, where models place stopping probability on different EOS tokens. Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigate…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. LastOPD: Taming Collapse in Latent On-Policy Distillation

    The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

    Hugging Face Daily Papersarxiv.org2 minpaper