proomt

Search

Search posts, papers, and topics

All posts

Hugging FaceEsakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov, Sagar Davasam9 min readintermediate

AutoSynthData: Generating Training Data for Enterprise Agents

Summary

AutoSynthData generates synthetic training data for enterprise AI agents by identifying a target model's weaknesses and using a stronger teacher to guide task creation. It produces feasible, realistic, and difficult tasks, validated through execution and verification, to iteratively improve agent performance.

  • Training enterprise agents requires identifying specific capability gaps in their target environment.
  • Useful training tasks must be feasible, realistic, and challenging for the agent, defined by a system spec, user prompt, and verifier.
  • AutoSynthData uses a stronger "teacher" model to characterize successful behavior for identified capability gaps.
  • Generated tasks undergo positive and negative verification, and a repair loop, to ensure high quality and correctness.

This system is crucial for engineers building enterprise AI agents, as it provides a structured, automated way to generate high-quality, targeted training data for specific capability gaps in complex environments.

7/10

Related reading

  1. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

    Hugging Face Daily Papersarxiv.org2 minpaper
  2. SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation

    SkillGym is an automated pipeline that generates verifiable environments and training data to improve LLM agents' ability to use skills for complex tasks. It constructs 6.8k environments and 19k trajectories, demonstrating that finetuning significantly boosts LLM performance and skill invocation rates across various models and benchmarks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  3. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  4. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

    The paper introduces BI‑Bench, a new benchmark of real‑world BI questions derived from public dashboards, and BI‑Agent, a tool‑augmented LLM system that breaks BI workflows into search, join, and transform subtasks. Baseline LLMs hit <50 % accuracy on BI‑Bench. By orchestrating specialized data‑management tools and post‑training the model with supervised fine‑tuning and reinforcement learning on…

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. Agentic coding in the enterprise: Is your pipeline ready?

    Agentic coding lets AI agents write, test, and submit code autonomously, shifting the bottleneck from writing to governing code in production. Enterprises face rising failures, unclear ownership, growing costs, and weakened controls, which require a unified pipeline visibility layer.

    Codeshipcloudbees.com6 min