proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersShiyang Zhou, Xionghao Wu, Wenbo Li1 min readpaperadvanced

EVO-WAM: Evolving World Action Models through Video-Action Verification

Summary

EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.

  • EVO-WAM adds state prediction and anchored multi-frame context to WAMs, enabling fully autoregressive rollouts without external feedback.
  • Task‑completing prefixes are selected by a vision‑language model and filtered for video‑action consistency using an inverse dynamics model.
  • Iterative training on verified prefixes improves simulated task success up to 2.5× and real‑world composite task success by 56.7 percentage points.
  • The approach eliminates the need to execute candidate actions in the environment, reducing reliance on additional demonstrations.

Robotics engineers focusing on imitation learning and policy adaptation should care because it offers a practical way to improve policies for new tasks without extra demonstrations or environment interaction.

8/10

Related reading

  1. Native Action-Prior Learning from Videos for World Action Models

    Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

    WorldLine is an action-driven visual simulator for robotic manipulation that decouples transferable dynamics learning from heterogeneous action grounding. It improves prediction accuracy and task success by learning from vast action-free and action-conditioned robot video data, enabling more efficient policy evaluation and embodied planning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

    Dream4ACT introduces a shared visual action interface (“action views”) that renders robot joint configurations from four virtual cameras using URDF forward kinematics, letting a single video autoencoder and diffusion transformer model both observations and actions across different robot embodiments. Trained with masked flow‑matching, the model attains 88.98% success on RoboTwin 2.0 and a 65.66 ov…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Distill frozen world‑model features into a Vision‑Language‑Action policy via a single alignment loss; no teacher at train time, no extra runtime cost. 0.8 B student runs 32 ms / 1.86 GB on RTX 5090, hits 97.9 % on LIBERO and improves RoboCasa‑GR1 from 48.2 % to 50.5 %, with transfer to real single‑arm and bimanual robots.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper