proomt

Search

Search posts, papers, and topics

multi teacher

RSS
  1. 2

    PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

    The paper presents PMOPD, a projection-based method that tracks low-dimensional subspaces of task-specific parameter updates during multi-teacher on-policy distillation and removes interfering components. Experiments on Qwen2.5-7B and Llama-3.1-8B show consistent 2-point gains across code, reasoning, and math tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. 3

    Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

    On‑policy distillation of LLMs traditionally uses KL divergence, but this paper shows that merely aligning the update direction toward the teacher—via a simple (+1/‑1) token‑wise reward—achieves the same effect. Building on this, they introduce Consensus Multi‑Teacher OPD, which lets every sample learn from all teachers and consistently beats the prior single‑teacher approach on math and code ben…

    Hugging Face Daily Papersarxiv.org2 minpaper