1
PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection
PANORAMA introduces a VLM that grounds each phrase of a caption to pixel masks by selecting from a phrase‑conditioned pool of mask proposals, trained jointly with caption generation. It sets a new benchmark on the PanoCaps dataset and outperforms prior methods on several pixel‑level grounding tasks.
Hugging Face Daily Papersarxiv.org1 minpaper
