Hugging Face Daily PapersRan Dan, Si-Tong Wei, Pengfei Xiong1 min readpaperadvanced
Octrees as an Explicit 3D Language
Summary
OctLLM treats 3D geometry as a sequence of octree occupancy tokens, using a Sparse Octree to keep sequences short while preserving shape. It adds lightweight 3D branches to a frozen vision‑language backbone, achieving state‑of‑the‑art image‑to‑3D generation with far fewer trainable parameters.
- OctLLM encodes geometry as explicit octree occupancy tokens, preserving spatial structure for LLMs.
- Sparse Octree (S‑Octree) randomly drops penultimate nodes, reducing sequence length without losing shape fidelity.
- 3D capacity is added via independent trainable branches in selected transformer blocks, keeping the pretrained backbone frozen.
- Achieves 17.4% lower image‑to‑3D FID and 28.7‑point improvement in render‑grounded captioning over ShapeLLM‑Omni.
Teams building multimodal LLMs for 3D tasks need a way to retain spatial detail while keeping training costs low.
7/10
