Hugging Face Daily PapersYiqi Liu, Ruifeng Yuan, Yang Wang1 min readpaperadvanced
World Embedding Benchmark
Summary
This paper introduces the World Embedding Benchmark (WEB), a dataset of 8,000 simulated physics videos with annotations, to evaluate how video embeddings encode physical information. It finds current models struggle with physical alignment and reveals a trade-off between cross-modal alignment and quantitative property recoverability.
- WEB provides 8,000 simulated physics videos across 80 families with physical annotations for benchmarking.
- Pre-trained omnimodal models perform poorly on physical text-video retrieval and classification tasks.
- Lightweight probes can extract quantitative physical properties from frozen video embeddings.
- Continual contrastive training improves retrieval but degrades quantitative property regression, showing a trade-off.
Researchers and engineers working on physically-aware world models and video generation should care, as this benchmark provides a critical tool for evaluating and improving physical fidelity.
8/10