Hugging Face Daily PapersZhaochong An, Fei Zhang, Menglin Jia1 min readpaperadvanced
Native Action-Prior Learning from Videos for World Action Models
Summary
Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.
- NAVA‑WAM pretrains action policies directly from observation‑only videos using future‑video flow‑matching and joint attention.
- Training proceeds in two stages: video‑only pretraining then fine‑tuning with a small set of action‑labeled demos via joint video‑action flow matching.
- Asymmetric attention decouples visual processing from iterative action denoising, enabling efficient action‑only inference.
- Experiments report consistent outperformance over prior methods on in‑distribution and out‑of‑distribution tasks with fewer action labels.
Robotics engineers who want to reduce reliance on costly action‑annotated trajectories should care.
6/10