proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYiqi Liu, Ruifeng Yuan, Yang Wang1 min readpaperadvanced

World Embedding Benchmark

Summary

This paper introduces the World Embedding Benchmark (WEB), a dataset of 8,000 simulated physics videos with annotations, to evaluate how video embeddings encode physical information. It finds current models struggle with physical alignment and reveals a trade-off between cross-modal alignment and quantitative property recoverability.

  • WEB provides 8,000 simulated physics videos across 80 families with physical annotations for benchmarking.
  • Pre-trained omnimodal models perform poorly on physical text-video retrieval and classification tasks.
  • Lightweight probes can extract quantitative physical properties from frozen video embeddings.
  • Continual contrastive training improves retrieval but degrades quantitative property regression, showing a trade-off.

Researchers and engineers working on physically-aware world models and video generation should care, as this benchmark provides a critical tool for evaluating and improving physical fidelity.

8/10

Related reading

  1. ROWBench: Do Video Models Render What the Program Specifies?

    PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

    The paper investigates why video diffusion models often break physical laws, pinpointing excessive spatial decay from Rotary Position Embedding (RoPE) as the culprit. By analyzing cross‑attention trajectories and self‑attention patterns, the authors identify specific attention heads that drive motion planning. They propose a lightweight fix: scaling RoPE frequency per denoising step, which empiri…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

    APM-Bench is a new benchmark for evaluating persistent memory in egocentric streaming video assistants across intermittent sessions. It reveals a significant utility-latency-storage trade-off, showing current models struggle with long-term recall, low overhead, and proactive assistance simultaneously.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    OSWorld-Science introduces a benchmark suite of 146 scientific software tasks for evaluating visual language model agents across domains like molecular design and image analysis. The authors provide a harness for systematic comparison and show that current state‑of‑the‑art VLMs still struggle with many scientific workflows.

    Hugging Face Daily Papersarxiv.org2 minpaper