proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersPengyu Zhu, Jingyi Yang, Yi Liu, Li Sun, Sen Su1 min readpaperadvanced

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Summary

SkillDRE is an automated framework that evolves malicious AI agent skills using a dual-stage feedback loop, combining pre-execution scanning and runtime defense feedback. It achieved a 45.28% attack success rate against victim models, significantly outperforming baselines while bypassing scanners and maintaining benign functionality.

  • SkillDRE automates red-teaming for AI agent skills by evolving malicious payloads.
  • It uses a dual-stage feedback loop: scanner-guided pre-execution and runtime-guided refinement.
  • Malicious skills are evolved to bypass defenses while preserving benign task capability.
  • Achieved 45.28% attack success rate, exceeding the strongest baseline by 40.3%.

AI safety and security engineers should care as this framework demonstrates a robust method for evolving evasive malicious AI agent skills, exposing potential vulnerabilities in current defense strategies.

8/10

Related reading

  1. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  2. SkillSpec: Intent-Masked Specification Reasoning for Agent Skill Correctness

    SkillSpec introduces a Hoare‑style framework that turns heterogeneous agent skill artifacts into a unified graph and reasons about correctness via intent‑masked specifications. In a study of 515 real‑world skills it flagged 763 confirmed defects with 61.2% precision, especially exposing intent‑implementation mismatches.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization

    COBRA‑Skills uses a contextual‑bandit loop to selectively evaluate and evolve LLM agent skills, achieving better performance with roughly half the evaluation cost of prior methods. The framework works with limited examples and remains robust across different agent setups and even when the target model creates its own skills.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  5. Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

    Code2Skill is an automated pipeline that mines popular GitHub repositories to extract verifiable, implementation‑anchored procedural “skills”. It builds a bank of ~1 M skill records (atomic ops, workflows, patterns) with provenance metadata, verifies each via blind reconstruction, and shows that augmenting LLM‑based agents with these skills yields an average 11.7% performance lift across 72 proto…

    Hugging Face Daily Papersarxiv.org1 minpaper
  6. Skill2Real: Agentic Skill Learning for Zero-Shot Sim-to-Real Robot Manipulation

    Skill2Real is an agentic framework that learns robot manipulation skills in simulation via a Proposer‑Verifier‑Governor loop and a hierarchical Cerebellum‑Brain memory, then transfers them zero‑shot to real robots using a shared API. Experiments on LIBERO‑90 and Robosuite show up to ~79% success on real tasks without any real‑world fine‑tuning, and ablations confirm the Verifier and Governor are…

    Hugging Face Daily Papersarxiv.org1 minpaper