proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXin Li, Hao Jiang, Xin Gao1 min readpaperadvanced

Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation

Summary

The paper shows that standard multi-teacher on-policy distillation (MOPD) gives unbalanced updates, hurting specialist knowledge like math. By measuring each domain’s feedback spread and normalizing it (DN-MOPD), they rebalance updates, recover most math performance, and improve average scores on six benchmarks.

  • MOPD’s student underperforms the best single specialist because instruction-following feedback dominates due to higher variance.
  • DN-MOPD rescales each domain's token-level feedback by its measured spread, equalizing influence across domains.
  • Experiments on Qwen3.5 models (three sizes) across six benchmarks show consistent average score gains over vanilla MOPD.
  • Ablations reveal the improvement mainly stems from down-weighting instruction-following feedback rather than boosting math feedback.

Engineers merging specialist LLMs need a principled way to balance their contributions, not just routing, to retain domain-specific strengths.

7/10

Related reading

  1. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

    On‑policy distillation of LLMs traditionally uses KL divergence, but this paper shows that merely aligning the update direction toward the teacher—via a simple (+1/‑1) token‑wise reward—achieves the same effect. Building on this, they introduce Consensus Multi‑Teacher OPD, which lets every sample learn from all teachers and consistently beats the prior single‑teacher approach on math and code ben…

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

    The paper presents PMOPD, a projection-based method that tracks low-dimensional subspaces of task-specific parameter updates during multi-teacher on-policy distillation and removes interfering components. Experiments on Qwen2.5-7B and Llama-3.1-8B show consistent 2-point gains across code, reasoning, and math tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper