proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersAhmadreza Jeddi, Enming Zhang, Jasper Gerigk1 min readpaperadvanced

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

Summary

Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

  • VLM inference is costly due to processing images/videos as long sequences of visual tokens.
  • Performance degradation from token pruning is partly due to underutilization of remaining information, termed the "representation-utilization gap".
  • SCOPD uses on-policy self-distillation with a full-context teacher to train a student on pruned visual tokens.
  • SCOPD+ enhances this by selectively distilling visually sensitive response positions with a small visual-budget intervention.

Engineers and researchers deploying vision-language models will find this method useful for significantly reducing inference costs while maintaining high performance.

8/10

Related reading

  1. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Where‑OPD introduces on‑policy self‑distillation for multimodal LLMs where the teacher receives spatially grounded textual guidance from synthetic scenes. Post‑training on these annotation‑free scenes improves real‑world vision‑language benchmarks by an average of 3.23 points.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. In-Context Robot Learning with VLM Agents

    GPT‑Policy is a framework that lets a large vision‑language model (e.g. GPT‑6 Astra) perform in‑context robot learning: a context compiler extracts visual transitions from demos, the VLM proposes tool actions, and a constrained controller verifies and executes them. Real‑robot experiments show that raw video demos improve success rates even without explicit action labels, and that providing align…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Distill frozen world‑model features into a Vision‑Language‑Action policy via a single alignment loss; no teacher at train time, no extra runtime cost. 0.8 B student runs 32 ms / 1.86 GB on RTX 5090, hits 97.9 % on LIBERO and improves RoboCasa‑GR1 from 48.2 % to 50.5 %, with transfer to real single‑arm and bimanual robots.

    Hugging Face Daily Papersarxiv.org1 minpaper