proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersBingxuan Li, Siqi Song, Yizhuo Wu1 min readpaperadvanced

MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation

Summary

MotorMind is a robot manipulation harness that enables general-purpose Vision Language Models (VLMs) to perform zero-shot robot manipulation without task-specific training or external tools. It achieves significantly higher success rates on benchmark suites and real robots compared to prior zero-shot methods.

  • MotorMind connects VLM-proposed mid-level actions to deterministic robot control and feedback.
  • It operates without task-specific policy training, coding agents, or additional grounding tools.
  • Achieves 66.7% success on LIBERO-PRO suites and 53.8% under perturbations, outperforming prior methods.
  • Demonstrates 95% average success on a real xArm6 robot across various settings.

Robotics engineers and researchers can leverage this work to explore more generalized and adaptable robot manipulation systems using off-the-shelf VLMs, reducing the need for extensive specialized training.

8/10

Related reading

  1. Transferring the Intelligence of VLMs to Robotic Control

    RoboDawn lets a vision‑language model (VLM) drive a robot via a tiny discrete command set (translate/rotate/gripper). Using a few in‑context demos, the VLM learns the interface and task strategy, then runs closed‑loop: observe image → reason → act → re‑observe. On the RoboTwin 2.0 C2R benchmark RoboDawn hits 53.2 % success zero‑shot, 73.6 % with one demo (vs. 46 % baseline). On RoboDojo it goes f…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. In-Context Robot Learning with VLM Agents

    GPT‑Policy is a framework that lets a large vision‑language model (e.g. GPT‑6 Astra) perform in‑context robot learning: a context compiler extracts visual transitions from demos, the VLM proposes tool actions, and a constrained controller verifies and executes them. Real‑robot experiments show that raw video demos improve success rates even without explicit action labels, and that providing align…

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

    Skill2Real is an agentic framework that learns robot manipulation skills in simulation via a Proposer‑Verifier‑Governor loop and a hierarchical Cerebellum‑Brain memory, then transfers them zero‑shot to real robots using a shared API. Experiments on LIBERO‑90 and Robosuite show up to ~79% success on real tasks without any real‑world fine‑tuning, and ablations confirm the Verifier and Governor are…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. WorldLine: Action-Driven Visual Simulation for Robotic Manipulation

    WorldLine is an action-driven visual simulator for robotic manipulation that decouples transferable dynamics learning from heterogeneous action grounding. It improves prediction accuracy and task success by learning from vast action-free and action-conditioned robot video data, enabling more efficient policy evaluation and embodied planning.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper