Hugging Face Daily PapersMohamed Eltahir, Fardows Adam, Duaa M. Tahir1 min readpaperadvanced
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
Summary
AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.
- AnswerMap generates spatial interpretability maps for VLMs using only the output head, making it black-box.
- It is training-free and task-agnostic, constructed by querying the VLM with image bands and yes/no questions.
- Validation shows AnswerMap is more faithful than attention maps, flipping 53% of correct answers upon region deletion vs 19% for attention.
- It enables native derivation of continuous outputs, bypassing reliance on discrete text tokens for tasks like localization.
This paper offers a novel, black-box approach to VLM interpretability and control, which is crucial for understanding and improving the reliability of these complex models in real-world applications.
8/10