Google DevelopersRavisri Valluri, Sagar Chapara, Rishabh Manoj11 min readadvanced
Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs
Summary
This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.
- Video diffusion attention scales quadratically with sequence length, becoming a major bottleneck at higher resolutions.
- Attention in video diffusion is often sparse, with distinct spatial and temporal head patterns that can be exploited.
- Naive sparse attention implementations can be slower than dense due to overheads like elementwise masking within tiles.
- Specializing tile execution with mask-free fast paths for "full" tiles significantly reduces latency on TPUs.
Engineers working on high-performance video generation models, especially on custom ML accelerators like TPUs, will find this a valuable case study in low-level optimization.
7/10

