Hugging Face Daily PapersHaotian Luo, Haoyu Wang, Zeyu Qin2 min readpaperadvanced
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
Summary
AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.
- AutoDataBench evaluates agent-generated tasks against production criteria: validity, novelty, difficulty, and behavioral coverage.
- Existing agents scored below 20/100 on AutoDataBench with a 45-minute time budget per task.
- While more time improved task quality, the cost per usable task remained high, indicating inefficiency.
- Automating data generation is a critical step for scaling LLM self-improvement loops beyond human labor.
This paper is important for researchers and engineers working on LLM self-improvement and autonomous agents, as it highlights a significant bottleneck in automated data generation and provides a new, practical evaluation framework.
7/10

