proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersXiangyu Zhu, Jin Xu, Yue Guo1 min readpaperadvanced

Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling

Summary

Dream4ACT introduces a shared visual action interface (“action views”) that renders robot joint configurations from four virtual cameras using URDF forward kinematics, letting a single video autoencoder and diffusion transformer model both observations and actions across different robot embodiments. Trained with masked flow‑matching, the model attains 88.98% success on RoboTwin 2.0 and a 65.66 ov…

  • Action views encode joint states as multi‑camera images, preserving embodiment geometry while unifying action representation.
  • A single diffusion‑based video model learns forward dynamics, inverse dynamics, and joint observation‑action generation by masking future frames.
  • Executable joint commands are recovered from predicted views using a URDF‑constrained, training‑free multiview reconstruction, avoiding per‑embodiment decoders.
  • Achieves 88.98% task success on RoboTwin 2.0 and 65.66 overall on TriWorldBench, showing effective closed‑loop manipulation across robots.

Engineers building multi‑robot manipulation systems need a unified, vision‑based action representation that works across embodiments and leverages video priors.

7/10

Related reading

  1. HuRo: Robotizing Human Videos for Scalable VLA Pretraining

    The paper introduces a pipeline that converts heterogeneous human videos into robot‑aligned observations and actions, creating the 630K‑episode HuRo dataset. Pretraining vision‑language‑action (VLA) policies on this data boosts real‑world manipulation success from ~51% to ~80% and improves out‑of‑distribution robustness.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. EVO-WAM: Evolving World Action Models through Video-Action Verification

    EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

    Ego2Act introduces a new benchmark for evaluating goal-directed manipulation in egocentric video generation, featuring 2,640 videos across 110 real-world tasks. It reveals that current video generation models struggle with multi-step physical reasoning, often skipping steps and failing at fine-grained object manipulation and persistent world modeling.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Native Action-Prior Learning from Videos for World Action Models

    Native Action-Prior Learning (NAVA‑WAM) trains robot action policies directly from observation‑only videos by matching future video flow through a joint attention mechanism, then fine‑tunes with a small set of labeled demos. Experiments show it beats prior methods on both in‑distribution and out‑of‑distribution tasks and transfers to real robots with fewer action labels.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    Action tokenizers for autoregressive VLA models often fail to preserve subtle action adjustments, despite good pointwise reconstruction. This paper introduces Physical Rank Consistency (PRC) to measure relational fidelity and ActionPiece, a new tokenizer that uses joint supervision to preserve these physical relationships, significantly improving policy success on robotics benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper