proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersPatrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana1 min readpaperadvanced

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Summary

Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

  • Ego2Act is a new benchmark for evaluating egocentric video generation models on multi-step, goal-directed manipulation tasks.
  • The benchmark comprises 2,640 videos from 110 real-world tasks, varying in object clutter and multi-step complexity.
  • Ego2ActJudge is a reference-free evaluation pipeline for task completion and physics plausibility, aligning well with human judgment.
  • Current video generation models frequently skip or partially execute steps, leading to unfulfilled high-level goals.

This benchmark is crucial for researchers developing video generation models and embodied AI, providing a rigorous testbed to advance physically plausible, goal-directed simulation capabilities.

8/10

Related reading

  1. Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

    Real2Gym is an agentic Real2Sim2Real framework that converts real-world human and robot videos into interactive simulation environments for robot skill learning. It reconstructs scenes, validates actions, and distills skills, achieving higher success rates than GPT-6 Astra in both simulation and real robot tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

    WorldLine is an action-driven visual simulator for robotic manipulation that decouples transferable dynamics learning from heterogeneous action grounding. It improves prediction accuracy and task success by learning from vast action-free and action-conditioned robot video data, enabling more efficient policy evaluation and embodied planning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

    Skill2Real is an agentic framework that learns robot manipulation skills in simulation via a Proposer‑Verifier‑Governor loop and a hierarchical Cerebellum‑Brain memory, then transfers them zero‑shot to real robots using a shared API. Experiments on LIBERO‑90 and Robosuite show up to ~79% success on real tasks without any real‑world fine‑tuning, and ablations confirm the Verifier and Governor are…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

    Dream4ACT introduces a shared visual action interface (“action views”) that renders robot joint configurations from four virtual cameras using URDF forward kinematics, letting a single video autoencoder and diffusion transformer model both observations and actions across different robot embodiments. Trained with masked flow‑matching, the model attains 88.98% success on RoboTwin 2.0 and a 65.66 ov…

    Hugging Face Daily Papersarxiv.org1 minpaper