proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersYuhan Guo, Jinming Liu, Liang Xu1 min readpaperadvanced

EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making

Summary

LLM agents struggle with generalization to new environments. EVOKE is a post-training method that elicits an agent's pretrained world knowledge by supervising action preferences under diverse goals at fixed states, improving transferability without explicit world model training.

  • LLM agents often rely on superficial contextual habits, hindering generalization to unseen environments.
  • EVOKE introduces goal diversity at fixed states during post-training to force agents to leverage internal world knowledge.
  • The method supervises action preferences for the same state under *different* goals, making superficial policies fail.
  • EVOKE improves task performance, generalization to unseen environments, and data efficiency across diverse tasks.

This method is important for researchers and practitioners building LLM agents, as it offers a novel, data-efficient way to improve agent generalization and robustness in diverse environments.

8/10

Related reading

  1. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. EVO-WAM: Evolving World Action Models through Video-Action Verification

    EVO-WAM adapts world action models to unseen robot tasks by generating video-action rollouts and self‑verifying them with vision-language and inverse dynamics models, removing the need for external execution or new demonstrations. It raises simulated RoboTwin 2.0 success rates from ~27% to 68% (Cosmos3) and ~28% to 46% (DreamZero), and boosts real‑world composite task success from 20% to 76.7%.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning

    RetireOPD introduces a self‑retiring on‑policy distillation framework for multi‑turn RL agents. A skill‑conditioned teacher is first trained with environment rewards, then a skill‑free student learns jointly via RL and token‑level distillation. The student automatically drops the teacher once its performance gap stops shrinking and it reaches a target success‑rate fraction, after which training c…

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

    COBRA‑Skills uses a contextual‑bandit loop to selectively evaluate and evolve LLM agent skills, achieving better performance with roughly half the evaluation cost of prior methods. The framework works with limited examples and remains robust across different agent setups and even when the target model creates its own skills.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

    EvoSkill‑GUI lets GUI agents revise their procedural skills on‑the‑fly without extra training by using a reflect‑revise‑reuse loop that edits skill packages during execution. The approach yields up to +16.2% improvement on MobileWorld and similar gains on AndroidWorld and OSWorld, and the evolved skills transfer to related tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper