proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXirui Li, Peng Shi, Mingwen Dong2 min readpaperadvanced

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Summary

LEGO-Anything is an Image-to-Code framework where a coding agent iteratively writes and executes Blender code to reconstruct 3D scenes from single images. It introduces LEGO-Bench for evaluation, showing GPT-6-astra performs best but struggles with scene initialization and self-evaluation, which LEGO-Plugin helps mitigate.

  • LEGO-Anything reconstructs 3D scenes as editable Blender code, not fixed 3D models.
  • LEGO-Bench is a new simulator-grounded benchmark for evaluating 3D scene recovery.
  • GPT-6-astra achieved 53.4% indoor and 39.6% outdoor scores on LEGO-Bench.
  • Key agent issues include weak initialization, regressive edits, and unreliable self-evaluation.

This work is significant for researchers and engineers developing AI agents for complex creative tasks and 3D scene understanding, offering a new paradigm for editable 3D reconstruction.

8/10

Related reading

  1. SNAP3D: Physically Grounded 3D Parts for Assembly from a Single Image

    The paper presents SNAP3D, a physics‑guided pipeline that converts a single image into a set of 3D parts that can be assembled without interpenetration. By using simulation‑driven connector placement and a new physics‑based evaluation, the method yields assemblies that are both geometrically accurate and stable enough for 3D printing.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Octrees as an Explicit 3D Language

    OctLLM treats 3D geometry as a sequence of octree occupancy tokens, using a Sparse Octree to keep sequences short while preserving shape. It adds lightweight 3D branches to a frozen vision‑language backbone, achieving state‑of‑the‑art image‑to‑3D generation with far fewer trainable parameters.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Towards Self-Driving Codebases

    The post argues that AI agents could eventually handle low‑level engineering tasks—bug fixing, debugging, UI consistency, growth experiments—if the dev toolchain is made “agent‑legible”. It outlines missing primitives (global memory, code‑base rot prevention, better dev environments) and proposes a bootstrapping process to measure and improve a repo’s “agent readiness”. The piece is largely specu…

    Hacker News front pagedetail.dev9 minHN12099
  4. CodeMidas: Scaling Agentic Coding RL Environments from Code Itself

    CodeMidas builds RL environments directly from open‑source code: agents explore a repo, infer a spec, generate tests from the original implementation, and filter tasks via execution checks. The pipeline yields 5,545 high‑quality coding tasks across 23 languages and 15 domains. Training the MiMo‑V2.5 agent with GRPO on this dataset improves benchmark scores by 8‑18% (e.g., DeepSWE +11.7%, ProgramB…

    Hugging Face Daily Papersarxiv.org1 minpaper