proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersGuillaume Besset, Erwann Carn, Timothée Carecchio1 min readpaperadvanced

OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport

Summary

The paper presents OTRetarget, a method that jointly retargets human motion and object trajectories to humanoid robots using entropic optimal transport to map surface interaction descriptors. On the OMOMO benchmark it reaches 87 % Jaccard similarity and 8.7 mm depth error, and is validated on a real G1 robot with RL policies.

  • Entropic optimal transport transfers signed‑distance and direction interaction features from human to robot/object meshes.
  • A constrained inverse‑kinematics formulation jointly optimizes robot and object poses per frame while preserving contacts.
  • Achieves 87 % Jaccard interaction score and 8.7 mm depth error on OMOMO, far surpassing OmniRetarget.
  • Demonstrated on a physical G1 humanoid using whole‑body policies trained via reinforcement learning on retargeted references.

Robotics engineers building humanoid manipulation pipelines need a principled way to transfer human demonstrations while keeping contact fidelity.

7/10

Related reading

  1. Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

    This paper introduces Movement Trend Guidance (MTG), a method to provide foresight to 3D diffusion policies for robotic manipulation without explicit trajectory planning. MTG learns a compact latent representation of interaction evolution, significantly improving performance on various benchmarks with minimal parameter overhead.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite

    This paper introduces Recursive Self-Rewrite (RSR), a framework that enables a single base LLM to discover solutions for complex tasks using diverse specialized environments (harnesses). It then reconstructs these successful trajectories into training data suitable for a general environment, significantly improving the model's performance on various benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Transferring the Intelligence of VLMs to Robotic Control

    RoboDawn lets a vision‑language model (VLM) drive a robot via a tiny discrete command set (translate/rotate/gripper). Using a few in‑context demos, the VLM learns the interface and task strategy, then runs closed‑loop: observe image → reason → act → re‑observe. On the RoboTwin 2.0 C2R benchmark RoboDawn hits 53.2 % success zero‑shot, 73.6 % with one demo (vs. 46 % baseline). On RoboDojo it goes f…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. EvolvingAvatar: Interactive 3D Head Generation That Adapts as Conversations Unfold

    EvolvingAvatar is a novel 3D head generation model that adapts to ongoing conversations using test-time training and a self-supervised dyadic context prediction objective. It improves conversational motion statistics by learning from user audiovisual input during interaction, reducing expression mismatch by up to 11.1%.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Fingers as Legs: Learning Self-Supported Locomotion and Manipulation with an Anthropomorphic Hand

    The authors train an anthropomorphic robotic hand to crawl, steer, and recover from falls using its fingers for both support and manipulation, via a reinforcement‑learning reward formulation tuned to the hand's asymmetry. Sim‑to‑real experiments show faster locomotion than quadruped‑style rewards and successful untethered tasks without onboard vision.

    Hugging Face Daily Papersarxiv.org1 minpaper