proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXinge Peng, Yiting Lu, Tianwu Zhi1 min readpaperadvanced

PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop

Summary

This paper introduces PhysVista, a new benchmark to evaluate the physical intelligence of Vision-Language Models (VLMs) using a closed perception-reasoning-assessment loop. Experiments reveal significant limitations in current VLMs' physical reasoning and plausibility assessment capabilities.

  • PhysVista evaluates VLM physical intelligence via a human-inspired perception-reasoning-assessment loop.
  • It jointly assesses physical state perception, dynamics reasoning, and plausibility assessment.
  • The benchmark distinguishes between event-level and scale-level physical reasoning for fine-grained analysis.
  • It incorporates both real-world and AI-generated videos to cover diverse scenarios.

Researchers and developers working on VLMs should care about this benchmark as it provides a structured way to diagnose and improve physically grounded multimodal intelligence.

7/10

Related reading

  1. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control

    VA‑Bench is a new benchmark that evaluates general‑purpose multimodal LLMs on the full observe‑reason‑act‑revise loop in embodied robotics, using RGB demonstrations, active camera control, and metric Cartesian commands. The best model reaches 53.9% average task success, showing active perception helps but long‑horizon tasks remain unsolved.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model

    This paper introduces a novel evaluation framework to assess the physical world reasoning capabilities of omni-modal generative models like MiniMax-H3. It found that MiniMax-H3 achieved an overall success rate of 41.97% across 517 instances, with significant performance variations depending on the input modalities and reasoning tasks.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision

    This paper evaluates GPT-6 Astra and five other frontier general-purpose AI systems across 34 computer vision capabilities and 55 benchmarks. It finds these systems excel at semantic interpretation and reasoning, but struggle with metric geometric accuracy, faithful reconstruction, and fine-grained specialized knowledge.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper