Hugging Face Daily PapersLanglin Huang, Hao Liu, Mononito Goswami1 min readpaperadvanced
On the Off-Policy Teacher in On-Policy Distillation
Summary
On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.
- OPD teachers degrade when supervising student-generated, off-policy trajectories.
- Teacher continuation performance worsens as student-generated prefixes lengthen.
- SCOUT is a co-training framework that adapts the teacher to student prefixes.
- SCOUT uses RL with verifiable rewards to optimize the teacher's conditional ability.
Engineers working on improving large language model efficiency and performance through distillation will find this relevant for addressing a core limitation in current OPD methods.
8/10