proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersTongbo Chen, Junbo Niu, Zhengxi Lu1 min readpaperadvanced

HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents

Summary

The paper introduces HybridCUA, a framework that trains computer-use agents to interleave GUI actions with command‑line commands. Using a new 5 K hybrid trajectory dataset and CLI‑aware RL rewards, the 9 B‑parameter model improves OSWorld accuracy by 14.8 pts and shows gains on WindowsAgentArena.

  • HybridCUA dataset contains 5K hybrid GUI/CLI trajectories and 3K verified RLVR tasks.
  • Training combines supervised fine‑tuning on hybrid trajectories with RL using CLI‑aware rewards.
  • HybridCUA‑9B reaches 53.6% accuracy on OSWorld, a 14.8‑point gain over the base model.
  • Performance also improves by 4.0 points on WindowsAgentArena, showing cross‑platform benefits.

Developers of autonomous desktop agents can adopt the hybrid GUI/CLI approach to boost efficiency and reliability across platforms.

7/10

Related reading

  1. RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents

    The paper presents RecreationWorld, a five‑platform framework that lets hybrid computer‑use agents learn by recreating the behavior of a running reference, and introduces RecreationBench, a 250‑task benchmark with programmatic and visual assertions. Experiments show GPT‑6 Astra reaches 58.1% overall but struggles with deeper programmatic tests, highlighting gaps in current agents.

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents

    EvoSkill‑GUI lets GUI agents revise their procedural skills on‑the‑fly without extra training by using a reflect‑revise‑reuse loop that edits skill packages during execution. The approach yields up to +16.2% improvement on MobileWorld and similar gains on AndroidWorld and OSWorld, and the evolved skills transfer to related tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

    OSWorld-Science introduces a benchmark suite of 146 scientific software tasks for evaluating visual language model agents across domains like molecular design and image analysis. The authors provide a harness for systematic comparison and show that current state‑of‑the‑art VLMs still struggle with many scientific workflows.

    Hugging Face Daily Papersarxiv.org2 minpaper
  4. CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents

    CERA-MoA proposes a reinforcement‑learning loop where a router and a set of LLM agents are trained together. A “familiarity” estimator reads mid‑layer hidden states to predict each agent’s competence on a query, letting the router activate only a minimal subset of agents that meet a cumulative confidence threshold. The system also feeds targeted training examples to agents based on their evolving…

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Build Agentic UI with the new Blazor AI components

    The new experimental Blazor AI components provide building blocks for "Agentic UI," enabling developers to integrate AI agent interactions into Blazor applications. They offer components and a state model to render streamed agent output, tool calls, and shared application state.

    .NETmicrosoft.com11 min