proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersShenghe Zheng, Wenbo Li, Jiyao Zhang1 min readpaperadvanced

WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

Summary

WorldLine is an action-driven visual simulator for robotic manipulation that decouples transferable dynamics learning from heterogeneous action grounding. It improves prediction accuracy and task success by learning from vast action-free and action-conditioned robot video data, enabling more efficient policy evaluation and embodied planning.

  • WorldLine decouples dynamics learning from action grounding, enabling transferability across robot embodiments.
  • It uses an image-space action representation for a shared control interface across diverse robots.
  • Trained on over 10,000 hours of action-free and 2,000 hours of action-conditioned robot videos.
  • Achieves 0.1626 higher robot-mask IoU on failed trajectories and 74% trajectory success prediction.

Robotics engineers and researchers can use WorldLine for more scalable and efficient policy evaluation and embodied planning, reducing the need for costly real-world data collection.

8/10

Related reading

  1. EVO-WAM: Evolving World Action Models through Video-Action Verification

    EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. In-Context Robot Learning with VLM Agents

    GPT‑Policy is a framework that lets a large vision‑language model (e.g. GPT‑6 Astra) perform in‑context robot learning: a context compiler extracts visual transitions from demos, the VLM proposes tool actions, and a constrained controller verifies and executes them. Real‑robot experiments show that raw video demos improve success rates even without explicit action labels, and that providing align…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Native Action-Prior Learning from Videos for World Action Models

    Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.

    Hugging Face Daily Papersarxiv.org1 minpaper