Hugging Face Daily PapersXu Wan, Wenyue Xu, Shengjie Zhao1 min readpaperadvanced
Mitigating the Length-Scaling Tax with Online Distillation
Summary
The authors define the length‑scaling tax (LST) as excess response length without accuracy gain and propose Length Self‑Distillation (LSD), an online EMA‑based teacher that requires no external model. Experiments show LSD matches or exceeds RL performance while cutting LST from 19% to -3.7% on single‑turn and from 31.4% to 13.7% on multi‑turn tasks.
- LST quantifies unnecessary verbosity in RL‑fine‑tuned LLMs.
- LSD uses an exponential moving average of the current policy as a teacher, avoiding extra models.
- LSD retains the original RL objective for unsolved prompts, focusing distillation on solved ones.
- Empirically, LSD reduces LST by up to ~23% on single‑turn and ~17% on multi‑turn queries.
LLM engineers fine‑tuning with reinforcement learning should care because LSD curtails unnecessary response length without sacrificing task performance.
8/10