proomt

Search

Search posts, papers, and topics

All posts

Google DevelopersRavisri Valluri, Sagar Chapara, Rishabh Manoj11 min readadvanced

Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

Summary

This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.

  • Video diffusion attention scales quadratically with sequence length, becoming a major bottleneck at higher resolutions.
  • Attention in video diffusion is often sparse, with distinct spatial and temporal head patterns that can be exploited.
  • Naive sparse attention implementations can be slower than dense due to overheads like elementwise masking within tiles.
  • Specializing tile execution with mask-free fast paths for "full" tiles significantly reduces latency on TPUs.

Engineers working on high-performance video generation models, especially on custom ML accelerators like TPUs, will find this a valuable case study in low-level optimization.

7/10

Related reading

  1. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention introduces a training‑free low‑bit attention pipeline for diffusion transformers. It smooths value tensors via lightweight online clustering (V‑Smooth) and quantizes only the residual after subtracting block means, restoring the mean from the softmax row sum. It also replaces the FP32 softmax exponential with a fused FP8 cast (ExpCast‑FP8) that maps log‑scores directly to E4M3 probab…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  5. AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

    The paper presents AV‑GRPO, a diffusion‑based reinforcement‑learning framework that treats audio and video generation as separate but coordinated tasks by anchoring rollouts to each modality and freezing the opposite tower during optimization. Evaluated on the new 5DAV dataset and benchmarks, it achieves higher fidelity, better text‑modality alignment, and tighter audio‑video sync than the previo…

    Hugging Face Daily Papersarxiv.org1 minpaper