1
PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
This paper introduces PhysVista, a new benchmark to evaluate the physical intelligence of Vision-Language Models (VLMs) using a closed perception-reasoning-assessment loop. Experiments reveal significant limitations in current VLMs' physical reasoning and plausibility assessment capabilities.
Hugging Face Daily Papersarxiv.org1 minpaper
