Hugging Face Daily PapersXinge Peng, Yiting Lu, Tianwu Zhi1 min readpaperadvanced
PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
Summary
This paper introduces PhysVista, a new benchmark to evaluate the physical intelligence of Vision-Language Models (VLMs) using a closed perception-reasoning-assessment loop. Experiments reveal significant limitations in current VLMs' physical reasoning and plausibility assessment capabilities.
- PhysVista evaluates VLM physical intelligence via a human-inspired perception-reasoning-assessment loop.
- It jointly assesses physical state perception, dynamics reasoning, and plausibility assessment.
- The benchmark distinguishes between event-level and scale-level physical reasoning for fine-grained analysis.
- It incorporates both real-world and AI-generated videos to cover diverse scenarios.
Researchers and developers working on VLMs should care about this benchmark as it provides a structured way to diagnose and improve physically grounded multimodal intelligence.
7/10