proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersZihan Wang, Zhen Wu, Pieter Abbeel1 min readpaperadvanced

Counterfactual Video Generation Enables Scalable Humanoid Loco-Manipulation

Summary

PRISM is a real-to-sim-to-real framework that uses counterfactual video generation to create diverse training data for humanoid loco-manipulation. It enables training a single policy that generalizes to unseen objects without real-world fine-tuning, addressing data scarcity in robotics.

  • PRISM generates diverse "counterfactual" human-object interaction videos from a few real exemplar videos.
  • A contact-anchored real-to-sim pipeline reconstructs human and object motions into physically plausible trajectories.
  • The intra-class variability from generated videos allows training a single policy that generalizes to unseen objects.
  • The policy deploys on a real robot using only onboard depth observations, requiring no real-world fine-tuning.

This work is significant for robotics engineers and researchers as it offers a scalable solution to the data collection bottleneck for training complex humanoid loco-manipulation skills.

8/10

Related reading

  1. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand

    The authors train an anthropomorphic robotic hand to crawl, steer, and recover from falls using its fingers for both support and manipulation, via a reinforcement‑learning reward formulation tuned to the hand's asymmetry. Sim‑to‑real experiments show faster locomotion than quadruped‑style rewards and successful untethered tasks without onboard vision.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

    Real2Gym is an agentic Real2Sim2Real framework that converts real-world human and robot videos into interactive simulation environments for robot skill learning. It reconstructs scenes, validates actions, and distills skills, achieving higher success rates than GPT-6 Astra in both simulation and real robot tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper