Hugging Face Daily PapersYouzhi Liu, Ruobing Zheng, Boyuan Tong1 min readpaperadvanced
PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
Summary
The paper presents PMOPD, a projection-based method that tracks low-dimensional subspaces of task-specific parameter updates during multi-teacher on-policy distillation and removes interfering components. Experiments on Qwen2.5-7B and Llama-3.1-8B show consistent 2-point gains across code, reasoning, and math tasks.
- Task-specific parameter updates rapidly concentrate in low-dimensional subspaces, enabling subspace memory construction.
- PMOPD projects gradients and optimizer updates to eliminate components that conflict with protected task directions.
- A lightweight conflict probe guides task ordering, while a cycling strategy balances subspace estimation and task revisitation.
- Average scores improve by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B across code, reasoning, and math benchmarks.
LLM engineers and researchers need robust techniques to combine multiple specialized capabilities without cross-task interference during distillation.
7/10