proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXincheng He, Siyu Ma, Chang Yu1 min readpaperadvanced

Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

Summary

Skill2Real is an agentic framework that learns robot manipulation skills in simulation via a Proposer‑Verifier‑Governor loop and a hierarchical Cerebellum‑Brain memory, then transfers them zero‑shot to real robots using a shared API. Experiments on LIBERO‑90 and Robosuite show up to ~79% success on real tasks without any real‑world fine‑tuning, and ablations confirm the Verifier and Governor are…

  • PVG loop uses privileged simulation evidence to verify skill updates while keeping skills expressed through a public robot API.
  • Hierarchical memory separates low‑level manipulation (Cerebellum) from high‑level task composition (Brain); only the Brain is trained on tasks.
  • Zero‑shot transfer achieves ~79% mean completion on four real‑world manipulation tasks without any real‑world fine‑tuning.
  • Ablations show removing the Verifier or Governor drops success by 17.3 and 13.3 percentage points respectively.

Robotics engineers seeking high‑performing sim‑to‑real transfer without costly real‑world data will find a concrete architecture that delivers strong zero‑shot results.

8/10

Related reading

  1. Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

    Real2Gym is an agentic Real2Sim2Real framework that converts real-world human and robot videos into interactive simulation environments for robot skill learning. It reconstructs scenes, validates actions, and distills skills, achieving higher success rates than GPT-6 Astra in both simulation and real robot tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

    Code2Skill is an automated pipeline that mines popular GitHub repositories to extract verifiable, implementation‑anchored procedural “skills”. It builds a bank of ~1 M skill records (atomic ops, workflows, patterns) with provenance metadata, verifies each via blind reconstruction, and shows that augmenting LLM‑based agents with these skills yields an average 11.7% performance lift across 72 proto…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

    SkillDRE is an automated framework that evolves malicious AI agent skills using a dual-stage feedback loop, combining pre-execution scanning and runtime defense feedback. It achieved a 45.28% attack success rate against victim models, significantly outperforming baselines while bypassing scanners and maintaining benign functionality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

    The paper presents PARTS, a framework that augments a frozen pretrained robot policy with RL‑learned residuals on selected bottleneck subtasks, using local success rewards and minimal human resets. In real‑world bimanual and single‑arm tasks, PARTS more than doubles success rates with only minutes of robot rollouts, outperforming prior fine‑tuning methods.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper