proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersHaotian Luo, Haoyu Wang, Zeyu Qin2 min readpaperadvanced

AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

Summary

AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

  • AutoDataBench evaluates agent-generated tasks against production criteria: validity, novelty, difficulty, and behavioral coverage.
  • Existing agents scored below 20/100 on AutoDataBench with a 45-minute time budget per task.
  • While more time improved task quality, the cost per usable task remained high, indicating inefficiency.
  • Automating data generation is a critical step for scaling LLM self-improvement loops beyond human labor.

This paper is important for researchers and engineers working on LLM self-improvement and autonomous agents, as it highlights a significant bottleneck in automated data generation and provides a new, practical evaluation framework.

7/10

Related reading

  1. AutoSynthData: Generating Training Data for Enterprise Agents

    AutoSynthData generates synthetic training data for enterprise AI agents by identifying a target model's weaknesses and using a stronger teacher to guide task creation. It produces feasible, realistic, and difficult tasks, validated through execution and verification, to iteratively improve agent performance.

    Hugging Facehuggingface.co9 min
  2. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  3. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

    The paper introduces BI‑Bench, a new benchmark of real‑world BI questions derived from public dashboards, and BI‑Agent, a tool‑augmented LLM system that breaks BI workflows into search, join, and transform subtasks. Baseline LLMs hit <50 % accuracy on BI‑Bench. By orchestrating specialized data‑management tools and post‑training the model with supervised fine‑tuning and reinforcement learning on…

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. Last 3 days: AI Evals, October cohort

    This post announces a live, hands-on course on building reliable evaluation systems for production AI agents, with enrollment closing soon. The course covers designing evals for quality, safety, cost, and latency, and building LLM-as-a-Judge systems.

    ByteByteGobytebytego.com1 minrelease