proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersJunlan Xiao, Junwei Jiang, Zaibin Zhang1 min readpaperintermediate

CARE: Experience-Guided Atomic Corrective Execution for Vision-Language-Action Policies

Summary

CARE learns corrective actions from real failure rollouts of Vision‑Language‑Action policies, models stage‑conditioned deviation distributions, and synthesizes representative failure states plus corrective demos. At inference it uses stage‑wise planning plus 3‑D monitoring to trigger atomic adjustments, improving task‑success rates by ~15 pts in simulation and real‑world dual‑arm tasks. Introduce…

  • Collect failed rollouts instead of hand‑crafted perturbations to capture realistic failure modes.
  • Model post‑failure deviations conditioned on the current manipulation stage, yielding empirical distributions of failure states.
  • Sample these distributions to generate synthetic failure states and corresponding corrective demonstrations for training.
  • During execution, combine stage‑wise planner with 3‑D perception monitoring to detect deviations and apply atomic corrective actions or re‑operations while preserving progress.

Robotic manipulation systems that rely on Vision‑Language‑Action policies are fragile when the robot deviates from the nominal trajectory, leading to costly failures in real deployments. CARE’s data‑driven failure modeling and atomic correction mechanism directly addresses this brittleness, offerin…

7/10

Related reading

  1. In-Context Robot Learning with VLM Agents

    GPT‑Policy is a framework that lets a large vision‑language model (e.g. GPT‑6 Astra) perform in‑context robot learning: a context compiler extracts visual transitions from demos, the VLM proposes tool actions, and a constrained controller verifies and executes them. Real‑robot experiments show that raw video demos improve success rates even without explicit action labels, and that providing align…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models

    Latent Interface Training (LIT) first learns a goal‑conditioned action prior without visual input, then adds a pose‑supervised latent interface as the only visual conditioning path. Applied to several vision‑language‑action models, LIT cuts vision‑action shortcuts and lifts LIBERO‑Plus success by 3.9–10.7 points and real‑world task success by 13.3–16.7 points under distribution shifts.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

    OmniHarness introduces a symbolic‑policy framework that extracts reusable procedural knowledge from multimodal LLM‑driven visual generation runs. By decoupling task logic from instance inputs, the system can instantiate, adapt, and compose policies for new visual tasks, using intermediate verification for on‑the‑fly refinement while keeping the underlying model frozen. Self‑directed practice task…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models

    Action tokenizers for autoregressive VLA models often fail to preserve subtle action adjustments, despite good pointwise reconstruction. This paper introduces Physical Rank Consistency (PRC) to measure relational fidelity and ActionPiece, a new tokenizer that uses joint supervision to preserve these physical relationships, significantly improving policy success on robotics benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. From Pretraining to Proficiency: Real-World Subtask RL for Long-Horizon Manipulation with Minimal Human Intervention

    The paper presents PARTS, a framework that augments a frozen pretrained robot policy with RL‑learned residuals on selected bottleneck subtasks, using local success rewards and minimal human resets. In real‑world bimanual and single‑arm tasks, PARTS more than doubles success rates with only minutes of robot rollouts, outperforming prior fine‑tuning methods.

    Hugging Face Daily Papersarxiv.org1 minpaper