Hugging Face Daily PapersSeo Hyun Kim, Sunwoo Hong, Younwoo Choi1 min readpaperadvanced
Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Summary
Pivot-SD is an efficient self-distillation framework for masked diffusion language models (dLMs) that addresses their credit-assignment problem by training only on high-impact "pivot" tokens. It improves LLaDA-8B-Instruct performance on math and code benchmarks using minimal data compared to other methods.
- dLMs face a credit-assignment challenge where early token commitments significantly shape the final output.
- Pivot-SD identifies "pivots" using an information-gain metric to measure uncertainty reduction over masked positions.
- It applies cross-entropy to pivots from successful trajectories and targeted unlikelihood to those from failed ones.
- The method is data-efficient, requiring only 200 questions and four rollouts per question for training.
Engineers working on training and fine-tuning diffusion language models can use this method to achieve better performance with significantly less data and computational cost.
8/10
