Hugging Face Daily PapersXin Li, Hao Jiang, Xin Gao1 min readpaperadvanced
Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Summary
The paper shows that standard multi-teacher on-policy distillation (MOPD) gives unbalanced updates, hurting specialist knowledge like math. By measuring each domain’s feedback spread and normalizing it (DN-MOPD), they rebalance updates, recover most math performance, and improve average scores on six benchmarks.
- MOPD’s student underperforms the best single specialist because instruction-following feedback dominates due to higher variance.
- DN-MOPD rescales each domain's token-level feedback by its measured spread, equalizing influence across domains.
- Experiments on Qwen3.5 models (three sizes) across six benchmarks show consistent average score gains over vanilla MOPD.
- Ablations reveal the improvement mainly stems from down-weighting instruction-following feedback rather than boosting math feedback.
Engineers merging specialist LLMs need a principled way to balance their contributions, not just routing, to retain domain-specific strengths.
7/10