proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersTan-Dzung Do, Tuan Dat Phuong, Nico Bohlinger1 min readpaperadvanced

CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments

Summary

CrossBFM distills a shared latent behavior space for humanoids, enabling a single vector to represent motion, pose, or reward across different robot embodiments. It drastically cuts training time from hundreds to ~11 GPU-hours and allows cross-embodiment transfer and generalization to unseen robots.

  • CrossBFM creates a unified latent behavior space transferable across various humanoid embodiments.
  • Training time is reduced from hundreds of GPU-hours per robot to ~11 GPU-hours for multiple robots.
  • The unified encoder architecture uses no robot-specific parameters, enabling simultaneous training.
  • It supports motion tracking, goal reaching, and reward optimization across distilled humanoids.

This work is crucial for robotics engineers and researchers, as it accelerates the development of generalizable humanoid control policies and reduces computational costs for training.

8/10

Related reading

  1. Grounded Action Model: 3D Grounding as a Foundation for Robotics

    The Grounded Action Model (GAM) adds explicit 3D metric grounding to robot foundation models via a shared object‑centric representation, improving robustness to scene changes. Experiments on RoboTwin 2.0, LIBERO‑PRO, and real robots show state‑of‑the‑art success rates, especially under visual shift and long‑horizon tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

    HarnessVLN introduces a zero‑shot, training‑free embodied navigation framework that wraps a multimodal LLM in an "Agent Harness" – a tool‑based protocol that validates planner actions against spatial evidence, tracks progress with hierarchical event memory, and maintains a persistent spatiotemporal graph for recovery. The system works for instruction‑following and object‑goal tasks, achieving 60.…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

    Dream4ACT introduces a shared visual action interface (“action views”) that renders robot joint configurations from four virtual cameras using URDF forward kinematics, letting a single video autoencoder and diffusion transformer model both observations and actions across different robot embodiments. Trained with masked flow‑matching, the model attains 88.98% success on RoboTwin 2.0 and a 65.66 ov…

    Hugging Face Daily Papersarxiv.org1 minpaper