1
Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies
Distill frozen world‑model features into a Vision‑Language‑Action policy via a single alignment loss; no teacher at train time, no extra runtime cost. 0.8 B student runs 32 ms / 1.86 GB on RTX 5090, hits 97.9 % on LIBERO and improves RoboCasa‑GR1 from 48.2 % to 50.5 %, with transfer to real single‑arm and bimanual robots.
Hugging Face Daily Papersarxiv.org1 minpaper
