1
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…
Hugging Face Daily Papersarxiv.org1 minpaper
