proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersHaoqiang Kang, Yinpeng Chen, Luyang Liu1 min readpaperadvanced

Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

Summary

The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

  • A dedicated scaffolding encoder replaces off‑the‑shelf vision encoders to produce task‑aligned latent visual tokens.
  • RL stage learns both mean and variance of the latent sampler, enabling stochastic exploration of latent trajectories.
  • Combined improvements raise FrozenLake spatial planning scores by +9.5 (up to +19 on 32×32 grids).
  • Average gain of +5.6 points across nine visual‑centric reasoning benchmarks demonstrates broader applicability.

Researchers and engineers building multimodal vision‑language models should care because better latent representations and stochastic RL sampling directly boost reasoning performance.

7/10

Related reading

  1. Reasoning with Image Generation

    ReImaGin proposes using image generation models as a flexible visual reasoning mechanism for multimodal LLMs, moving beyond rigid visual tools. It consistently outperforms text-only and specialist vision-tool baselines by up to 25% on diverse visual reasoning tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

    Reasoning vision-language models (VLMs) are expensive due to long visual token sequences. This paper identifies a "representation-utilization gap" where pruned visual tokens are underutilized, not just lost. SCOPD, a self-distillation framework, improves VLM performance with sparse visual contexts, achieving 92.43% of unpruned performance at 10% token retention.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper