1
AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.
Hugging Face Daily Papersarxiv.org1 minpaper
