Hugging Face Daily PapersXiangyu Zhu, Jin Xu, Yue Guo1 min readpaperadvanced
Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
Summary
Dream4ACT introduces a shared visual action interface (“action views”) that renders robot joint configurations from four virtual cameras using URDF forward kinematics, letting a single video autoencoder and diffusion transformer model both observations and actions across different robot embodiments. Trained with masked flow‑matching, the model attains 88.98% success on RoboTwin 2.0 and a 65.66 ov…
- Action views encode joint states as multi‑camera images, preserving embodiment geometry while unifying action representation.
- A single diffusion‑based video model learns forward dynamics, inverse dynamics, and joint observation‑action generation by masking future frames.
- Executable joint commands are recovered from predicted views using a URDF‑constrained, training‑free multiview reconstruction, avoiding per‑embodiment decoders.
- Achieves 88.98% task success on RoboTwin 2.0 and 65.66 overall on TriWorldBench, showing effective closed‑loop manipulation across robots.
Engineers building multi‑robot manipulation systems need a unified, vision‑based action representation that works across embodiments and leverages video priors.
7/10