proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersRenxi Wang, Mingshan Hee, Fajri Koto1 min readpaperadvanced

SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

Summary

SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

  • SkillGym automates the generation of verifiable environments and training data for LLM agent skill-use.
  • A builder-reviewer pipeline creates difficulty-controlled tasks with reference solutions and executable verifiers.
  • Finetuning LLMs (2B-122B params) with SkillGym's data improves performance on skill-use benchmarks.
  • Training increases agents' relevant skill invocation rate from 28% to 96%.

This work is important for researchers and engineers developing LLM agents, offering a systematic method to generate high-quality training data and improve agent reliability and performance on complex, skill-based tasks.

8/10

Related reading

  1. Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

    Code2Skill is an automated pipeline that mines popular GitHub repositories to extract verifiable, implementation‑anchored procedural “skills”. It builds a bank of ~1 M skill records (atomic ops, workflows, patterns) with provenance metadata, verifies each via blind reconstruction, and shows that augmenting LLM‑based agents with these skills yields an average 11.7% performance lift across 72 proto…

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness

    SkillSpec introduces a Hoare‑style framework that turns heterogeneous agent skill artifacts into a unified graph and reasons about correctness via intent‑masked specifications. In a study of 515 real‑world skills it flagged 763 confirmed defects with 61.2% precision, especially exposing intent‑implementation mismatches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

    SkillDRE is an automated framework that evolves malicious AI agent skills using a dual-stage feedback loop, combining pre-execution scanning and runtime defense feedback. It achieved a 45.28% attack success rate against victim models, significantly outperforming baselines while bypassing scanners and maintaining benign functionality.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

    Real2Gym is an agentic Real2Sim2Real framework that converts real-world human and robot videos into interactive simulation environments for robot skill learning. It reconstructs scenes, validates actions, and distills skills, achieving higher success rates than GPT-6 Astra in both simulation and real robot tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

    COBRA‑Skills uses a contextual‑bandit loop to selectively evaluate and evolve LLM agent skills, achieving better performance with roughly half the evaluation cost of prior methods. The framework works with limited examples and remains robust across different agent setups and even when the target model creates its own skills.

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

    EdgeGen automatically extracts compliance rules from a tool‑calling LLM agent’s specification and generates database‑grounded edge‑case tasks that violate those rules. Using these synthetic edge cases for finetuning and harness optimization improves benchmark performance by up to 42 % and 30 % respectively, without any human labeling.

    Hugging Face Daily Papersarxiv.org1 minpaper