proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersRuihan Yu, Yu-Ju Tsai, Muyao Niu1 min readpaperadvanced

EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling

Summary

EditHero is a new benchmark for long‑horizon, part‑level 3D editing that supplies natural‑language instructions, geometry and texture targets, and a deterministic engine that produces the exact result after each edit. Using it, the authors show that bottom‑up LLM/VLM agents preserve unchanged parts better than top‑down non‑agentic methods, though they run slower (minutes per edit).

  • EditHero provides sequences of part‑level 3D edits with NL instructions and target images for both geometry and texture.
  • A deterministic assembly engine can generate the exact target after every edit, and all sequences are hand‑reviewed for correctness.
  • Top‑down non‑agentic methods often modify regions that should stay fixed, while LLM/VLM agents follow instructions more faithfully.
  • LLM/VLM agents preserve unedited parts well but each edit takes minutes, highlighting a speed‑accuracy trade‑off.

Researchers building iterative 3D editing tools, especially those leveraging LLM/VLM agents, need realistic benchmarks to evaluate reliability and performance.

6/10

Related reading

  1. CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design

    CADWorld is a new benchmark suite of 200 long‑horizon mechanical CAD tasks in FreeCAD, covering sketching, part modeling, assembly, CAM, FEM, and more. Agents interact via screenshots and GUI actions; success is checked by executable validation of the saved CAD artifacts. Seven existing agents achieve at most 17.5 % success versus an 87 % expert baseline, highlighting the gap between GUI competen…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

    The paper presents GameHorizon, a unified suite comprising an automated annotation pipeline, a 5,000‑hour multi‑horizon gameplay dataset from 21 AAA titles, and reproducible offline and online benchmarks. Using it, the authors evaluate 47 models, exposing a hierarchy of task difficulty and gaps in long‑term planning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper