proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYugu Li, Zehong Cao, Peizhen Li1 min readpaperadvanced

Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

Summary

Existing self-distillation methods couple credit direction and magnitude, making them vulnerable to teacher errors and preference variance. Decoupled Credit Self-Distillation (DCSD) addresses this by theoretically separating credit direction and magnitude into two reliable signals, using belief-margin probing and marginal information gain to calibrate teacher supervision.

  • Existing self-distillation methods fundamentally couple credit direction and magnitude, leading to unreliable supervision.
  • DCSD decouples credit direction and magnitude, using belief-margin probing for direction and marginal information gain for magnitude.
  • This method calibrates privileged teacher supervision, enabling more reliable step-to-token credit assignment for policy optimization.
  • DCSD achieved superior performance across 11 benchmarks, improving mathematical reasoning by 8.45 points and multimodal reasoning by 7.01 points.

Machine learning researchers and practitioners working on self-distillation or policy optimization for large language models should care about this novel approach to improve model training and performance.

8/10

Related reading

  1. Mitigating the Length-Scaling Tax with Online Distillation

    The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. LastOPD: Taming Collapse in Latent On-Policy Distillation

    The authors find that latent supervision in on‑policy distillation can cause a dramatic performance collapse even as alignment improves. Their LastOPD method limits latent alignment to the final‑layer state and crossfades to token‑level supervision, delivering 4–5 point MATH‑500 gains and halving the steps to match token‑only OPD.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper