1
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Imagine3D-LLM augments a multimodal LLM with a small set of learnable summary tokens that are decoded into a compact 3D Gaussian splatting of the scene, supervised by photometric reconstruction. This auxiliary task improves cross‑view reasoning and yields better results on spatial‑reasoning benchmarks.
Hugging Face Daily Papersarxiv.org1 minpaper
