proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersSophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc1 min readpaperadvanced

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Summary

Where‑OPD introduces on‑policy self‑distillation for multimodal LLMs where the teacher receives spatially grounded textual guidance from synthetic scenes. Post‑training on these annotation‑free scenes improves real‑world vision‑language benchmarks by an average of 3.23 points.

  • Spatially grounded textual cues let a frozen teacher attend to relevant image regions during self‑distillation.
  • Procedurally generated scenes provide free object IDs and coordinates, enabling annotation‑free post‑training.
  • The student learns to reproduce teacher behavior from raw image+question, improving counting, document, and chart tasks.
  • Gains transfer to real benchmarks, yielding a 3.23‑point average boost across CVBench, V*, ZoomBench, BLINK, HR‑Bench, and MME‑RealWorld.

ML engineers building multimodal LLMs should care because it offers a cheap, annotation‑free way to boost visual reasoning performance.

7/10

Related reading

  1. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

    On‑policy distillation of LLMs traditionally uses KL divergence, but this paper shows that merely aligning the update direction toward the teacher—via a simple (+1/‑1) token‑wise reward—achieves the same effect. Building on this, they introduce Consensus Multi‑Teacher OPD, which lets every sample learn from all teachers and consistently beats the prior single‑teacher approach on math and code ben…

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. Region-Level Policy Optimization for Fine-grained MLLM Perception

    Vision‑RL2 trains a lightweight proposal network via region‑level reinforcement learning to select high‑resolution evidence for multimodal LLMs, allowing coarse‑resolution localization and fine‑resolution recognition. Across six fine‑grained vision benchmarks it reduces visual token count by ~4× while matching or surpassing full‑resolution accuracy.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

    OmniHarness introduces a symbolic‑policy framework that extracts reusable procedural knowledge from multimodal LLM‑driven visual generation runs. By decoupling task logic from instance inputs, the system can instantiate, adapt, and compose policies for new visual tasks, using intermediate verification for on‑the‑fly refinement while keeping the underlying model frozen. Self‑directed practice task…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. On the Off-Policy Teacher in On-Policy Distillation

    On-policy distillation (OPD) faces a challenge where the teacher's performance degrades when supervising student-generated, off-policy trajectories. The SCOUT framework addresses this by co-training the teacher to adapt to student-generated prefixes using reinforcement learning, consistently improving OPD effectiveness.

    Hugging Face Daily Papersarxiv.org1 minpaper