Hugging Face Daily PapersZhongbo Zhang, Jiayi Jin, Yifan Wang1 min readpaperadvanced
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control
Summary
VA‑Bench is a new benchmark that evaluates general‑purpose multimodal LLMs on the full observe‑reason‑act‑revise loop in embodied robotics, using RGB demonstrations, active camera control, and metric Cartesian commands. The best model reaches 53.9% average task success, showing active perception helps but long‑horizon tasks remain unsolved.
- Active camera control more than doubles success on a matched task (27.86% → 57.50%).
- Models achieve near‑perfect target localization (100%) but only ~79% on spatial relation inference.
- General‑purpose MLLMs struggle with long‑horizon multi‑object compositions, completing none fully.
- Held‑out geometry/layout variants can drop success by >30 points, highlighting limited transfer.
Anyone building embodied AI or robotic agents needs a realistic, open‑source benchmark to measure active perception and metric control capabilities of LLM‑based policies.
7/10