Hugging Face Daily PapersMikhail Dereviannykh, Vikram Voleti, Simon Donne1 min readpaperadvanced
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
Summary
The paper presents SemanTok, a flexible video tokenizer that injects frozen DINO features and reconstructs them from any token prefix, achieving strong semantic alignment and fidelity. A 201 M SemanTok AR model matches or exceeds a 3.4× larger VideoFlexTok baseline, with cheaper short‑prefix prediction and better generation quality.
- SemanTok feeds frozen DINO features into its encoder and adds lightweight heads that can reconstruct semantics from any retained token prefix.
- A 201 M SemanTok autoregressive model matches or beats a VideoFlexTok model 3.4× its size, demonstrating efficiency gains.
- Semantic alignment remains high on out‑of‑distribution classes and across all noise levels, even pure noise.
- Short token prefixes are cheaper to predict and improve generation fidelity, deferring pixel detail to later tokens.
Video generation researchers and engineers should care because SemanTok offers a more efficient tokenizer that preserves semantics, reducing model size and inference cost.
7/10
