proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSeo Hyun Kim, Sunwoo Hong, Younwoo Choi1 min readpaperadvanced

Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models

Summary

Pivot-SD is an efficient self-distillation framework for masked diffusion language models (dLMs) that addresses their credit-assignment problem by training only on high-impact "pivot" tokens. It improves LLaDA-8B-Instruct performance on math and code benchmarks using minimal data compared to other methods.

  • dLMs face a credit-assignment challenge where early token commitments significantly shape the final output.
  • Pivot-SD identifies "pivots" using an information-gain metric to measure uncertainty reduction over masked positions.
  • It applies cross-entropy to pivots from successful trajectories and targeted unlikelihood to those from failed ones.
  • The method is data-efficient, requiring only 200 questions and four rollouts per question for training.

Engineers working on training and fine-tuning diffusion language models can use this method to achieve better performance with significantly less data and computational cost.

8/10

Related reading

  1. Register Tokens for Bounded-State Reasoning in Diffusion Language Models

    Register tokens are fixed‑position embeddings that store a compact hidden state across diffusion‑based language model generation chunks, enabling bounded‑state reasoning without retaining all prior text. Post‑training on LLaDA and Dream shows up to +8.5 math and +19.5 code benchmark points versus plain text carry, and RL fine‑tuning further improves long‑horizon tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

    DACA‑GRPO adds denoising‑aware credit assignment to GRPO‑style RL trainers for diffusion LLMs. It computes per‑token importance scores from intermediate denoising steps and uses stratified masking to reduce mean‑field bias in likelihood estimates. Plug‑and‑play on three existing GRPO methods, it yields consistent gains on seven downstream tasks (up to +5.6 pp math, +7.4 pp code, +36.3 pp constrai…

    Apple Machine Learning Researchapple.com1 minpaper
  3. ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

    ROSS is a method for large language model post-training that selectively supervises historical self-generated rollouts, applying loss only to useful continuations while preserving full trajectory context. It consistently improves LLM performance across various tasks like code generation and instruction following, achieving gains on Qwen3.6-35B-A3B without requiring new policy rollouts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

    Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.

    Hugging Face Daily Papersarxiv.org1 minpaper