1
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?
AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.
Hugging Face Daily Papersarxiv.org2 minpaper
