proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersMohamed Eltahir, Fardows Adam, Duaa M. Tahir1 min readpaperadvanced

AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

Summary

AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.

  • AnswerMap generates spatial interpretability maps for VLMs using only the output head, making it black-box.
  • It is training-free and task-agnostic, constructed by querying the VLM with image bands and yes/no questions.
  • Validation shows AnswerMap is more faithful than attention maps, flipping 53% of correct answers upon region deletion vs 19% for attention.
  • It enables native derivation of continuous outputs, bypassing reliance on discrete text tokens for tasks like localization.

This paper offers a novel, black-box approach to VLM interpretability and control, which is crucial for understanding and improving the reliability of these complex models in real-world applications.

8/10

Related reading

  1. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper