proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZheng-Hui Huang, Guixu Lin, Yu-Ju Tsai1 min readpaperadvanced

ROWBench: Do Video Models Render What the Program Specifies?

Summary

PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.

  • PROWBench evaluates visual adherence of video models to explicit rules and program-specified world events.
  • It includes 170 episodes and 600 proxy videos, logging entity states and timestamped events for ground truth.
  • The benchmark introduces two VLM-based metrics: Logic-Render Alignment and Interaction Success Rate.
  • It provides an extensible framework for constructing scenes, controlling behaviors, and rendering multi-view observations.

Engineers and researchers developing programmable world models or next-generation game engines should care, as this benchmark offers a rigorous method to ensure visual outputs align with underlying program logic.

7/10

Related reading

  1. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    VA‑Bench is a new benchmark that evaluates general‑purpose multimodal LLMs on the full observe‑reason‑act‑revise loop in embodied robotics, using RGB demonstrations, active camera control, and metric Cartesian commands. The best model reaches 53.9% average task success, showing active perception helps but long‑horizon tasks remain unsolved.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. World Embedding Benchmark

    This paper introduces the World Embedding Benchmark (WEB), a dataset of 8,000 simulated physics videos with annotations, to evaluate how video embeddings encode physical information. It finds current models struggle with physical alignment and reveals a trade-off between cross-modal alignment and quantitative property recoverability.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

    OmniVBench is a new benchmark and the Omni‑R2V Dataset, offering 7 task families, 18 fine‑grained reference‑to‑video generation tasks and a factor‑grounded evaluation checklist of over 12 k items. The dataset provides 340 k industrial‑grade video samples and pipelines for constructing reference‑target pairs, exposing large performance gaps in current R2V models.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks

    ProgramDistill is a new benchmark that automatically extracts 1,975 replay‑verified feature behaviors from 26 real web apps, builds 4,063 coding‑agent tasks, and measures how well state‑of‑the‑art agents (e.g., GPT‑6 Astra, Claude Opus 5) can reconstruct full or partial applications, revealing steep drops in success as restoration depth grows.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

    CADWorld is a new benchmark suite of 200 long‑horizon mechanical CAD tasks in FreeCAD, covering sketching, part modeling, assembly, CAM, FEM, and more. Agents interact via screenshots and GUI actions; success is checked by executable validation of the saved CAD artifacts. Seven existing agents achieve at most 17.5 % success versus an 87 % expert baseline, highlighting the gap between GUI competen…

    Hugging Face Daily Papersarxiv.org1 minpaper