proomt

Search

Search posts, papers, and topics

All posts

Amazon ScienceStephen Zorio11 min readadvanced

A kernel-centric path to real-time video generation on Trainium

Summary

AWS Neuron Science and Reactor optimized real-time autoregressive diffusion video generation on Trainium. They achieved production-grade performance by using the Neuron Kernel Interface for direct hardware control and a hybrid sharding strategy to overcome challenges like dynamic shapes and memory access patterns.

  • Real-time autoregressive diffusion models for video generation face challenges from dynamic shapes, unusual memory access, and heavy cache management.
  • The Neuron Kernel Interface (NKI) provides direct hardware control to optimize bottlenecks, reducing 3D-RoPE kernel time from 5s to 1.8ms.
  • NKI-Dev-Suite generated optimized kernels, enabling the pipeline to use 11GB HBM where standard eager-mode ran out of memory.
  • A hybrid sharding strategy (sequence and tensor parallelism) was crucial for self-attention in long video sequences, which consumes 70% of compute.

Engineers deploying large generative AI models, especially for real-time video or interactive applications, will find this valuable for understanding hardware-specific optimization techniques.

7/10

Related reading

  1. PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

    PhysStream introduces a two‑stage autoregressive video generator that uses online‑derived positional and tracking maps (structured scene memory) and sparse velocity‑increment signals to enable fine‑grained, physics‑grounded control of multi‑object tabletop scenes. It cuts motion distribution error by 33 % and trajectory error by 12 % versus strong baselines, and wins 85 % of human preference test…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Streaming Video Editing with Easy Adaptation

    The paper introduces SVEET, a framework that adapts a pretrained bidirectional video diffusion model for streaming video editing via an auxiliary branch with temporally independent self‑attention and a decoupled orthogonal training scheme. It achieves real‑time 15 FPS editing on a single H100 GPU without extra acceleration.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

    Video DeltaNet (VDN) replaces full‑softmax attention in video diffusion models with a hybrid: per‑frame local Softmax for fine detail and a bidirectional linear memory (Video Delta Attention) for long‑range context. A teacher‑alignment schedule injects the linear branch into a pretrained MiniMax H3 model, preserving Softmax for text/audio streams. On eight NVIDIA B200 GPUs VDN‑H3 denoises a 14.3‑…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Accelerating Spatio-Temporal Attention for Video Diffusion on TPUs

    This post details how Google optimized spatio-temporal attention for video diffusion models on TPUs, turning theoretical sparsity into actual inference speedups. Key optimizations include specializing tile execution paths and carefully tuning tile sizes, resulting in significant latency reductions compared to dense attention.

    Google Developersgoogleblog.com11 min
  5. OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

    OmniVBench is a new benchmark and the Omni‑R2V Dataset, offering 7 task families, 18 fine‑grained reference‑to‑video generation tasks and a factor‑grounded evaluation checklist of over 12 k items. The dataset provides 340 k industrial‑grade video samples and pipelines for constructing reference‑target pairs, exposing large performance gaps in current R2V models.

    Hugging Face Daily Papersarxiv.org2 minpaper
  6. Video Generation Models: A Survey of Post-Training and Alignment

    This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

    Hugging Face Daily Papersarxiv.org1 minpaper