proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJaewoo Jung, Hyeonseo Yu, Honggyu An1 min readpaperadvanced

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Summary

Imagine3D-LLM augments a multimodal LLM with a small set of learnable summary tokens that are decoded into a compact 3D Gaussian splatting of the scene, supervised by photometric reconstruction. This auxiliary task improves cross‑view reasoning and yields better results on spatial‑reasoning benchmarks.

  • Adds learnable summary tokens after image patches and decodes them into a compact 3D Gaussian splatting representation.
  • Supervises the summary tokens with a photometric reconstruction loss while jointly training the next‑token prediction objective.
  • Reconstruction loss propagates 3D‑aware signals to image features, improving cross‑view correspondence without explicit geometry supervision.
  • Evaluated on multiple spatial‑reasoning and 3D understanding benchmarks, reporting consistent performance improvements over prior MLLM baselines.

ML engineers building vision‑language models that need to reason about 3D scenes will find a lightweight way to inject coarse geometry without full 3D supervision.

6/10

Related reading

  1. Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning

    The paper pinpoints two weaknesses in latent‑reasoning pipelines—suboptimal visual token encoders and deterministic RL sampling—and proposes Scaffolding Minds, which trains a dedicated scaffolding encoder and a stochastic RL sampler (learning both mean and variance). This yields up to +19 points improvement on spatial planning tasks and +5.6 points average gain on nine visual reasoning benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Reasoning with Image Generation

    ReImaGin proposes using image generation models as a flexible visual reasoning mechanism for multimodal LLMs, moving beyond rigid visual tools. It consistently outperforms text-only and specialist vision-tool baselines by up to 25% on diverse visual reasoning tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

    Where‑OPD introduces on‑policy self‑distillation for multimodal LLMs where the teacher receives spatially grounded textual guidance from synthetic scenes. Post‑training on these annotation‑free scenes improves real‑world vision‑language benchmarks by an average of 3.23 points.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors

    AnswerMap is a novel, training-free, black-box method for generating faithful spatial interpretability maps for Vision-Language Models (VLMs) directly from their output posteriors. It queries the VLM with image bands and yes/no relevance questions, demonstrating higher faithfulness than attention maps and enabling new applications like object localization and hallucination detection.

    Hugging Face Daily Papersarxiv.org1 minpaper