1
HuRo: Robotizing Human Videos for Scalable VLA Pretraining
The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.
Hugging Face Daily Papersarxiv.org1 minpaper
