proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersChaoyu Li, Xiaoyi Gu, Yogesh Kulkarni1 min readpaperadvanced

Video Generation Models: A Survey of Post-Training and Alignment

Summary

This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

  • Pretrained video models often lack reliable human intent following and temporal coherence.
  • Video alignment faces unique challenges like error accumulation and motion-appearance coupling.
  • Post-training is framed as a unifying framework for adapting pretrained models without retraining from scratch.
  • Approaches are categorized into supervised fine-tuning, self-training, preference-based, and inference-time methods.

Researchers and engineers working on controllable and reliable video generation will find this survey valuable for its structured conceptual foundation and practical guidance.

8/10

Related reading

  1. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  2. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

    The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. VideoGen-Agent: Reinforcing Video Generation Agents

    VideoGen-Agent is a multimodal RL‑trained agent that orchestrates augmentation, generation, and verification tools to improve text‑to‑video synthesis on a new 600‑prompt benchmark (VABench). It lifts a base generator’s score from 56.5 to 75.6 (‑19.1 pts) and to 86.1 when the generation tools are upgraded, with 84.3% human preference over the strongest baseline.

    Hugging Face Daily Papersarxiv.org1 minpaper