proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYouzhi Liu, Ruobing Zheng, Boyuan Tong1 min readpaperadvanced

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

Summary

The paper presents PMOPD, a projection-based method that tracks low-dimensional subspaces of task-specific parameter updates during multi-teacher on-policy distillation and removes interfering components. Experiments on Qwen2.5-7B and Llama-3.1-8B show consistent 2-point gains across code, reasoning, and math tasks.

  • Task-specific parameter updates rapidly concentrate in low-dimensional subspaces, enabling subspace memory construction.
  • PMOPD projects gradients and optimizer updates to eliminate components that conflict with protected task directions.
  • A lightweight conflict probe guides task ordering, while a cycling strategy balances subspace estimation and task revisitation.
  • Average scores improve by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B across code, reasoning, and math benchmarks.

LLM engineers and researchers need robust techniques to combine multiple specialized capabilities without cross-task interference during distillation.

7/10

Related reading

  1. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. LastOPD: Taming Collapse in Latent On-Policy Distillation

    The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper