Hugging Face Daily PapersBingxuan Li, Siqi Song, Yizhuo Wu1 min readpaperadvanced
MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
Summary
MotorMind is a robot manipulation harness that enables general-purpose Vision Language Models (VLMs) to perform zero-shot robot manipulation without task-specific training or external tools. It achieves significantly higher success rates on benchmark suites and real robots compared to prior zero-shot methods.
- MotorMind connects VLM-proposed mid-level actions to deterministic robot control and feedback.
- It operates without task-specific policy training, coding agents, or additional grounding tools.
- Achieves 66.7% success on LIBERO-PRO suites and 53.8% under perturbations, outperforming prior methods.
- Demonstrates 95% average success on a real xArm6 robot across various settings.
Robotics engineers and researchers can leverage this work to explore more generalized and adaptable robot manipulation systems using off-the-shelf VLMs, reducing the need for extensive specialized training.
8/10