Hugging Face Daily PapersZheng-Hui Huang, Guixu Lin, Yu-Ju Tsai1 min readpaperadvanced
ROWBench: Do Video Models Render What the Program Specifies?
Summary
PROWBench is a new benchmark designed to evaluate the visual fidelity of programmable world models to fine-grained, program-specified events and interactions. It comprises 170 programmatically constructed episodes and 600 proxy videos, using VLM-based metrics to check generated videos against observable consequences of program execution.
- PROWBench evaluates visual adherence of video models to explicit rules and program-specified world events.
- It includes 170 episodes and 600 proxy videos, logging entity states and timestamped events for ground truth.
- The benchmark introduces two VLM-based metrics: Logic-Render Alignment and Interaction Success Rate.
- It provides an extensible framework for constructing scenes, controlling behaviors, and rendering multi-view observations.
Engineers and researchers developing programmable world models or next-generation game engines should care, as this benchmark offers a rigorous method to ensure visual outputs align with underlying program logic.
7/10