Hugging Face Daily PapersChaoyu Li, Xiaoyi Gu, Yogesh Kulkarni1 min readpaperadvanced
Video Generation Models: A Survey of Post-Training and Alignment
Summary
This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.
- Pretrained video models often lack reliable human intent following and temporal coherence.
- Video alignment faces unique challenges like error accumulation and motion-appearance coupling.
- Post-training is framed as a unifying framework for adapting pretrained models without retraining from scratch.
- Approaches are categorized into supervised fine-tuning, self-training, preference-based, and inference-time methods.
Researchers and engineers working on controllable and reliable video generation will find this survey valuable for its structured conceptual foundation and practical guidance.
8/10
