proomt

Search

Search posts, papers, and topics

vlm

RSS
  1. 2

    AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

    AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. 3

    EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

    EditHero is a new benchmark for long‑horizon, part‑level 3D editing that supplies natural‑language instructions, geometry and texture targets, and a deterministic engine that produces the exact result after each edit. Using it, the authors show that bottom‑up LLM/VLM agents preserve unchanged parts better than top‑down non‑agentic methods, though they run slower (minutes per edit).

    Hugging Face Daily Papersarxiv.org1 minpaper