Hugging Face Daily PapersShiyang Zhou, Xionghao Wu, Wenbo Li1 min readpaperadvanced
EVO-WAM: Evolving World Action Models through Video-Action Verification
Summary
EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.
- EVO-WAM adds state prediction and anchored multi-frame context to WAMs, enabling fully autoregressive rollouts without external feedback.
- Task‑completing prefixes are selected by a vision‑language model and filtered for video‑action consistency using an inverse dynamics model.
- Iterative training on verified prefixes improves simulated task success up to 2.5× and real‑world composite task success by 56.7 percentage points.
- The approach eliminates the need to execute candidate actions in the environment, reducing reliance on additional demonstrations.
Robotics engineers focusing on imitation learning and policy adaptation should care because it offers a practical way to improve policies for new tasks without extra demonstrations or environment interaction.
8/10