proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersKerui Ren, Yingxiang Xu, Kaiwen Song1 min readpaperadvanced

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

Summary

Real2Gym is an agentic Real2Sim2Real framework that converts real-world human and robot videos into interactive simulation environments for robot skill learning. It reconstructs scenes, validates actions, and distills skills, achieving higher success rates than GPT-6 Astra in both simulation and real robot tasks.

  • Real2Gym is an agentic Real2Sim2Real framework for learning robot skills from video demonstrations.
  • It reconstructs editable 3D scenes from video, aligning objects and cameras for physics-based action validation.
  • Agents generate executable code for manipulation, observing outcomes to distill task procedures and recovery strategies.
  • Learned skills transfer to physical robots via a shared perception-control interface without model weight updates.

Robotics engineers and researchers can leverage this framework to efficiently acquire and transfer complex manipulation skills from diverse video demonstrations to physical robots.

8/10

Related reading

  1. Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

    Skill2Real is an agentic framework that learns robot manipulation skills in simulation via a Proposer‑Verifier‑Governor loop and a hierarchical Cerebellum‑Brain memory, then transfers them zero‑shot to real robots using a shared API. Experiments on LIBERO‑90 and Robosuite show up to ~79% success on real tasks without any real‑world fine‑tuning, and ablations confirm the Verifier and Governor are…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Grounded Action Model: 3D Grounding as a Foundation for Robotics

    The Grounded Action Model (GAM) adds explicit 3D metric grounding to robot foundation models via a shared object‑centric representation, improving robustness to scene changes. Experiments on RoboTwin 2.0, LIBERO‑PRO, and real robots show state‑of‑the‑art success rates, especially under visual shift and long‑horizon tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. EVO-WAM: Evolving World Action Models through Video-Action Verification

    EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

    Code2Skill is an automated pipeline that mines popular GitHub repositories to extract verifiable, implementation‑anchored procedural “skills”. It builds a bank of ~1 M skill records (atomic ops, workflows, patterns) with provenance metadata, verifies each via blind reconstruction, and shows that augmenting LLM‑based agents with these skills yields an average 11.7% performance lift across 72 proto…

    Hugging Face Daily Papersarxiv.org1 minpaper