Hugging Face Daily PapersSophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc1 min readpaperadvanced
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Summary
Where‑OPD introduces on‑policy self‑distillation for multimodal LLMs where the teacher receives spatially grounded textual guidance from synthetic scenes. Post‑training on these annotation‑free scenes improves real‑world vision‑language benchmarks by an average of 3.23 points.
- Spatially grounded textual cues let a frozen teacher attend to relevant image regions during self‑distillation.
- Procedurally generated scenes provide free object IDs and coordinates, enabling annotation‑free post‑training.
- The student learns to reproduce teacher behavior from raw image+question, improving counting, document, and chart tasks.
- Gains transfer to real benchmarks, yielding a 3.23‑point average boost across CVBench, V*, ZoomBench, BLINK, HR‑Bench, and MME‑RealWorld.
ML engineers building multimodal LLMs should care because it offers a cheap, annotation‑free way to boost visual reasoning performance.
7/10