proomt

Search

Search posts, papers, and topics

All posts

Hugging Face Daily PapersDeqing Fu, Huangyuan Su, Rajat Sen1 min readpaperadvanced

TabFM-Auto: Self-Evolving Pipelines for Tabular Foundation Models

Summary

TabFM-Auto couples a frozen tabular foundation model with an LLM agent that iteratively rewrites the data pipeline using metadata and validation signals. The approach lifts TabFM's Elo from 1785 to 2013 on TabArena and transfers to other models with +69‑+143 Elo gains.

  • An LLM agent can automatically refine cleaning, feature engineering, context selection, and post‑processing to improve a frozen tabular foundation model.
  • On the 51‑dataset TabArena benchmark, the best TabFM‑Auto configuration raises TabFM's Elo from 1785 to 2013.
  • Discovered pipelines transfer to other frozen tabular foundation models, adding 69–143 Elo without extra search.
  • TabFM‑Auto ranks first among MLE agents on the 8 tabular competitions of MLE‑Bench.

ML engineers building tabular pipelines or deploying foundation models should care because automated LLM‑driven pipeline evolution can deliver significant accuracy improvements with minimal human effort.

8/10

Related reading

  1. AutoSynthData: Generating Training Data for Enterprise Agents

    AutoSynthData generates synthetic training data for enterprise AI agents by identifying a target model's weaknesses and using a stronger teacher to guide task creation. It produces feasible, realistic, and difficult tasks, validated through execution and verification, to iteratively improve agent performance.

    Hugging Facehuggingface.co9 min
  2. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. AI for Games in the Foundation Model Era

    The paper surveys how foundation models are used across six roles in the game development lifecycle—from playing agents to design assistance and runtime adaptation. It highlights limited transferability due to game-specific interfaces and notes that evaluation is mature for bounded play but weak for adaptive and testing scenarios.

    Hugging Face Daily Papersarxiv.org1 minpaper
  4. Grounded Action Model: 3D Grounding as a Foundation for Robotics

    The Grounded Action Model (GAM) adds explicit 3D metric grounding to robot foundation models via a shared object‑centric representation, improving robustness to scene changes. Experiments on RoboTwin 2.0, LIBERO‑PRO, and real robots show state‑of‑the‑art success rates, especially under visual shift and long‑horizon tasks.

    Hugging Face Daily Papersarxiv.org1 minpaper
  5. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31