Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD is an efficient self-distillation framework for masked diffusion language models (dLMs) that addresses their credit-assignment problem by training only on high-impact "pivot" tokens. It improves LLaDA-8B-Instruct performance on math and code benchmarks using minimal data compared to other methods.
