Hugging Face Daily PapersJinho Jeong, Se June Joo, Jaehyun Kang1 min readpaperadvanced
HuRo: Robotizing Human Videos for Scalable VLA Pretraining
Summary
The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.
- Robotization pipeline infers missing intermediate signals to produce robot‑compatible observations and action trajectories from raw human videos.
- HuRo dataset contains ~630 K episodes and 142 M processed frames sourced from five public human‑video collections.
- Pretraining VLA policies on increasing amounts of HuRo data raises task completion from 51.5% to 80.3% on four real‑world manipulation tasks.
- Visual robotization steps improve out‑of‑distribution performance under spatial and visual shifts (34.9% → 72.2%).
Robotics researchers and engineers seeking scalable, low‑cost supervision for manipulation policies should care, as the work demonstrates that large‑scale human video can be turned into effective robot training data.
8/10
