Hugging Face Daily PapersPatrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana1 min readpaperadvanced
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation
Summary
Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.
- Ego2Act is a new benchmark for evaluating egocentric video generation models on multi-step, goal-directed manipulation tasks.
- The benchmark comprises 2,640 videos from 110 real-world tasks, varying in object clutter and multi-step complexity.
- Ego2ActJudge is a reference-free evaluation pipeline for task completion and physics plausibility, aligning well with human judgment.
- Current video generation models frequently skip or partially execute steps, leading to unfulfilled high-level goals.
This benchmark is crucial for researchers developing video generation models and embodied AI, providing a rigorous testbed to advance physically plausible, goal-directed simulation capabilities.
8/10