Hugging Face Daily PapersJulianna Piskorz, Antonin Berthon, Mihaela van der Schaar1 min readpaperadvanced
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
Summary
The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.
- Rollout policy (on/off) has limited effect on distillation performance compared to token‑level KL direction.
- Forward KL yields stable results regardless of rollout policy; reverse KL is sensitive and benefits from student‑generated rollouts.
- Learning rate is the dominant factor controlling catastrophic forgetting and update sparsity.
- On‑policy data can boost generalisation on harder arithmetic tasks, but the gain may vanish after subsequent RL fine‑tuning.
Engineers designing LLM distillation pipelines need to prioritize KL direction and learning‑rate tuning over rollout policy selection.
6/10