proomt

Search

Search posts, papers, and topics

All posts

Google ResearchYale Song and Yiwen Song, Research Scientists, Google10 min readadvanced

Automating coherent long-form video generation

Summary

Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

  • A hierarchical multi-agent framework (Co-Director) uses a Multi-Armed Bandit for global optimization of creative intent.
  • CANVAS ensures visual continuity in multi-shot narratives via persistent visual memory and structured representations of entities.
  • A²RD scales to minutes-long videos by adaptively switching between extrapolation and interpolation for temporal dynamics.
  • VQQA enables autonomous visual artifact refinement using VLM critiques as semantic gradients for closed-loop optimization.

Engineers building or integrating generative AI for complex, long-form content will find this framework valuable for its novel approach to maintaining consistency and coherence over time.

8/10

Related reading

  1. Video Generation Models: A Survey of Post-Training and Alignment

    This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. VideoGen-Agent: Reinforcing Video Generation Agents

    VideoGen-Agent is a multimodal RL‑trained agent that orchestrates augmentation, generation, and verification tools to improve text‑to‑video synthesis on a new 600‑prompt benchmark (VABench). It lifts a base generator’s score from 56.5 to 75.6 (‑19.1 pts) and to 86.1 when the generation tools are upgraded, with 84.3% human preference over the strongest baseline.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper