Hugging Face Daily PapersJaewoo Jung, Hyeonseo Yu, Honggyu An1 min readpaperadvanced
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Summary
Imagine3D-LLM augments a multimodal LLM with a small set of learnable summary tokens that are decoded into a compact 3D Gaussian splatting of the scene, supervised by photometric reconstruction. This auxiliary task improves cross‑view reasoning and yields better results on spatial‑reasoning benchmarks.
- Adds learnable summary tokens after image patches and decodes them into a compact 3D Gaussian splatting representation.
- Supervises the summary tokens with a photometric reconstruction loss while jointly training the next‑token prediction objective.
- Reconstruction loss propagates 3D‑aware signals to image features, improving cross‑view correspondence without explicit geometry supervision.
- Evaluated on multiple spatial‑reasoning and 3D understanding benchmarks, reporting consistent performance improvements over prior MLLM baselines.
ML engineers building vision‑language models that need to reason about 3D scenes will find a lightweight way to inject coarse geometry without full 3D supervision.
6/10