proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXu Wan, Wenyue Xu, Shengjie Zhao1 min readpaperadvanced

Mitigating the Length-Scaling Tax with Online Distillation

Summary

The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.

  • LST quantifies unnecessary verbosity in RL‑fine‑tuned LLMs.
  • LSD uses an exponential moving average of the current policy as a teacher, avoiding extra models.
  • LSD retains the original RL objective for unsolved prompts, focusing distillation on solved ones.
  • Empirically, LSD reduces LST by up to ~23% on single‑turn and ~17% on multi‑turn queries.

LLM engineers fine‑tuning with reinforcement learning should care because LSD curtails unnecessary response length without sacrificing task performance.

8/10

Related reading

  1. When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation

    On-policy distillation (OPD) can lead to excessively long student responses, a phenomenon called length inflation. This paper identifies "termination-token mismatch" between base students and post-trained teachers as a key source, where models place stopping probability on different EOS tokens. Treating functionally equivalent EOS tokens as a shared semantic stopping action substantially mitigate…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

    The paper systematically evaluates on‑policy vs off‑policy rollouts in LLM distillation, finding that rollout policy matters less than the KL direction and learning rate. Forward KL is robust to rollout choice, while reverse KL prefers student rollouts, and learning rate drives forgetting and sparsity.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Sharpening Tax in Post-Training

    The authors introduce Sharpening Tax, a metric that measures how much RL post‑training reduces test‑time scalability (pass@K) of LLM agents. They also propose Posterior‑Tempered Group Sampling, a simple temperature‑adaptation technique that lowers this tax and improves both single‑shot accuracy and coverage.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation

    Existing self-distillation methods couple credit direction and magnitude, making them vulnerable to teacher errors and preference variance. Decoupled Credit Self-Distillation (DCSD) addresses this by theoretically separating credit direction and magnitude into two reliable signals, using belief-margin probing and marginal information gain to calibrate teacher supervision.

    Hugging Face Daily Papersarxiv.org1 minpaper