proomt

Search

Search posts, papers, and topics

Hall of Fame

Hall of FameJonathan Ho, Ajay Jain, Pieter Abbeel202038 min readpaperadvanced

Denoising Diffusion Probabilistic Models

Summary

This paper demonstrates that Denoising Diffusion Probabilistic Models (DDPMs) can generate high-quality images, achieving state-of-the-art FID scores on CIFAR10. It establishes a novel connection between DDPMs and denoising score matching, leading to a simplified training objective that predicts the noise added at each step.

  • DDPMs achieve state-of-the-art image synthesis, outperforming prior generative models on metrics like FID.
  • A key contribution is the equivalence between DDPMs and denoising score matching with Langevin dynamics.
  • The simplified training objective involves a neural network learning to predict the noise component of a noisy image.
  • Sampling is a progressive denoising process, iteratively removing noise to reconstruct the original image.

This paper is foundational for modern diffusion models, providing the core theoretical and practical framework that enabled subsequent breakthroughs in generative AI.

9/10

Related reading

  1. DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

    DACA‑GRPO adds denoising‑aware credit assignment to GRPO‑style RL trainers for diffusion LLMs. It computes per‑token importance scores from intermediate denoising steps and uses stratified masking to reduce mean‑field bias in likelihood estimates. Plug‑and‑play on three existing GRPO methods, it yields consistent gains on seven downstream tasks (up to +5.6 pp math, +7.4 pp code, +36.3 pp constrai…

    Apple Machine Learning Researchapple.com1 minpaper
  2. Adversarial Training for Pixel Diffusion

    Pixel diffusion models often underrepresent fine-scale image statistics. This paper demonstrates that adversarial post-training, by adding an adversarial loss to non-high-noise timesteps, effectively corrects this deficiency. It restores missing high-frequency content, jointly improving distribution fidelity, coverage, prompt alignment, and perceptual quality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning

    The authors adapt Stable Diffusion 3 with a pruned text stream and patch‑wise normalization to condition on photogrammetric DSMs and Pléiades imagery, refining elevation maps. In French city tests the method halves RMSE, reaching 3.45 m in‑context and 2.77 m on a held‑out city.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs

    Flash-dLLM is a training-free framework that accelerates Diffusion LLM inference by addressing GPU memory I/O bottlenecks with an I/O-aware KV-cache kernel. It also introduces a KV-cache-driven draft-and-verify decoding strategy, achieving significant speedups (up to 11x) over prior methods.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention introduces a training‑free low‑bit attention pipeline for diffusion transformers. It smooths value tensors via lightweight online clustering (V‑Smooth) and quantizes only the residual after subtracting block means, restoring the mean from the softmax row sum. It also replaces the FP32 softmax exponential with a fused FP8 cast (ExpCast‑FP8) that maps log‑scores directly to E4M3 probab…

    Hugging Face Daily Papersarxiv.org1 minpaper