proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXin Lin, Zhifei Zhang, Yuqian Zhou1 min readpaperadvanced

Adversarial Training for Pixel Diffusion

Summary

Pixel diffusion models often underrepresent fine-scale image statistics. This paper demonstrates that adversarial post-training, by adding an adversarial loss to non-high-noise timesteps, effectively corrects this deficiency. It restores missing high-frequency content, jointly improving distribution fidelity, coverage, prompt alignment, and perceptual quality.

  • Adversarial post-training improves pixel diffusion models by adding an adversarial loss to predicted outputs at non-high-noise timesteps.
  • The method restores systematically underproduced high-frequency content in original pixel diffusion models.
  • It jointly enhances distribution fidelity, coverage, prompt alignment, and perceptual quality without altering model architecture or sampling.
  • Unlike perceptual loss, adversarial training increases high-frequency content without sacrificing distribution fidelity or prompt alignment.

Machine learning engineers working with image generation should care, as this method offers a significant improvement in the output quality and realism of pixel diffusion models.

8/10

Related reading

  1. Denoising Diffusion Probabilistic Models

    This paper demonstrates that Denoising Diffusion Probabilistic Models (DDPMs) can generate high-quality images, achieving state-of-the-art FID scores on CIFAR10. It establishes a novel connection between DDPMs and denoising score matching, leading to a simplified training objective that predicts the noise added at each step.

    Hall of Famearxiv.org38 minpaper
  2. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning

    The authors adapt Stable Diffusion 3 with a pruned text stream and patch‑wise normalization to condition on photogrammetric DSMs and Pléiades imagery, refining elevation maps. In French city tests the method halves RMSE, reaching 3.45 m in‑context and 2.77 m on a held‑out city.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

    This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.

    Google Developersgoogleblog.com11 min
  5. VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention introduces a training‑free low‑bit attention pipeline for diffusion transformers. It smooths value tensors via lightweight online clustering (V‑Smooth) and quantizes only the residual after subtracting block means, restoring the mean from the softmax row sum. It also replaces the FP32 softmax exponential with a fused FP8 cast (ExpCast‑FP8) that maps log‑scores directly to E4M3 probab…

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance

    FoMo introduces a fully automated method for generating perceptual distance labels between image pairs, eliminating the need for human annotation. It leverages the "forking moment" in a diffusion model's generative trajectory, where early divergence indicates large perceptual differences and late divergence indicates subtle ones.

    Hugging Face Daily Papersarxiv.org1 minpaper