Hugging Face Daily PapersWenxue Li, Peiyan Guan, Haoyang Jiang2 min readpaperadvanced
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation
Summary
OmniVBench is a new benchmark and the Omni‑R2V Dataset, offering 7 task families, 18 fine‑grained reference‑to‑video generation tasks and a factor‑grounded evaluation checklist of over 12 k items. The dataset provides 340 k industrial‑grade video samples and pipelines for constructing reference‑target pairs, exposing large performance gaps in current R2V models.
- OmniVBench defines 7 task families and 18 fine‑grained R2V tasks covering content, motion, style, structure, narrative, and multi‑reference control.
- It introduces a factor‑grounded evaluation with 12,172 checklist items to verify preservation, disentanglement, and routing of reference factors.
- The Omni‑R2V Dataset supplies 340 k processed video samples and scalable pipelines for building reference‑target pairs.
- Baseline tests show existing open‑ and closed‑source R2V models fall short across many tasks, highlighting current limitations.
Researchers and engineers building reference‑to‑video generation systems need comprehensive benchmarks and large‑scale training data to evaluate fine‑grained control capabilities.
6/10

