proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersRan Dan, Si-Tong Wei, Pengfei Xiong1 min readpaperadvanced

Octrees as an Explicit 3D Language

Summary

OctLLM treats 3D geometry as a sequence of octree occupancy tokens, using a Sparse Octree to keep sequences short while preserving shape. It adds lightweight 3D branches to a frozen vision‑language backbone, achieving state‑of‑the‑art image‑to‑3D generation with far fewer trainable parameters.

  • OctLLM encodes geometry as explicit octree occupancy tokens, preserving spatial structure for LLMs.
  • Sparse Octree (S‑Octree) randomly drops penultimate nodes, reducing sequence length without losing shape fidelity.
  • 3D capacity is added via independent trainable branches in selected transformer blocks, keeping the pretrained backbone frozen.
  • Achieves 17.4% lower image‑to‑3D FID and 28.7‑point improvement in render‑grounded captioning over ShapeLLM‑Omni.

Teams building multimodal LLMs for 3D tasks need a way to retain spatial detail while keeping training costs low.

7/10

Related reading

  1. LEGO-Anything: Coding Agents for 3D Scene Reconstruction

    LEGO-Anything is an Image-to-Code framework where a coding agent iteratively writes and executes Blender code to reconstruct 3D scenes from single images. It introduces LEGO-Bench for evaluation, showing GPT-6-astra performs best but struggles with scene initialization and self-evaluation, which LEGO-Plugin helps mitigate.

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation

    GEAR is a geometry‑enabled attention routing framework for long‑horizon camera‑controlled video generation. It treats per‑frame geometry as token‑level addresses, uses Geometric Correspondence Attention and an Invisible Octree to retrieve visual memory, achieving state‑of‑the‑art quality and consistent control over minute‑long videos.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Reasoning with Image Generation

    ReImaGin proposes using image generation models as a flexible visual reasoning mechanism for multimodal LLMs, moving beyond rigid visual tools. It consistently outperforms text-only and specialist vision-tool baselines by up to 25% on diverse visual reasoning tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene

    Mira-Scene proposes a compositional 3D scene reconstruction pipeline that replaces sparse pose regression with a dense Canonical Coordinate Map (CCM) linking image pixels to bounded object space. Coupled with a diffusion transformer, it achieves up to 40% higher 3D‑IoU than prior methods without scene‑level layout annotations.

    Hugging Face Daily Papersarxiv.org1 minpaper