proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersMikhail Dereviannykh, Vikram Voleti, Simon Donne1 min readpaperadvanced

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Summary

The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.

  • SemanTok feeds frozen DINO features into its encoder and adds lightweight heads that can reconstruct semantics from any retained token prefix.
  • A 201 M SemanTok autoregressive model matches or beats a VideoFlexTok model 3.4× its size, demonstrating efficiency gains.
  • Semantic alignment remains high on out‑of‑distribution classes and across all noise levels, even pure noise.
  • Short token prefixes are cheaper to predict and improve generation fidelity, deferring pixel detail to later tokens.

Video generation researchers and engineers should care because SemanTok offers a more efficient tokenizer that preserves semantics, reducing model size and inference cost.

7/10

Related reading

  1. Video Generation Models: A Survey of Post-Training and Alignment

    This survey reviews post-training and alignment strategies for video generation models, which often struggle with human intent and temporal coherence despite strong pretraining. It proposes a new taxonomy, categorizing methods into supervised fine-tuning, self-training, preference-based, and inference-time approaches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Automating coherent long-form video generation

    Google Research introduces a unified multi-agent framework to autonomously generate temporally consistent, long-form video narratives. It overcomes identity drift and cascading failures of current linear AI pipelines by treating generation as a global optimization and world-state tracking problem.

    Google Researchresearch.google10 min
  3. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

    OmniVBench is a new benchmark and the Omni‑R2V Dataset, offering 7 task families, 18 fine‑grained reference‑to‑video generation tasks and a factor‑grounded evaluation checklist of over 12 k items. The dataset provides 340 k industrial‑grade video samples and pipelines for constructing reference‑target pairs, exposing large performance gaps in current R2V models.

    Hugging Face Daily Papersarxiv.org2 minpaper