proomt

Search

Search posts, papers, and topics

All posts

TimescaleMatvey Arye8 min readintermediate

41 Experiments in 6 Days: What the Human Was Actually For

Summary

Timescale used an AI agent to run 41 experiments in six days, improving LLM conversational memory on the LoCoMo benchmark. The core insight is that humans are crucial for validating metrics and objectives, preventing the agent from optimizing against flawed measurements.

  • AI agents accelerate code and experiments, enabling 41 runs in 6 days for LLM memory development.
  • Humans are vital for validating metrics; agents will optimize flawed objectives without complaint.
  • Largest gains came from removing friction and improving data representation, not complex algorithms.
  • Using a single database (Postgres) for all search modes makes data representation changes cheap and fast.

Engineers working on LLM development or MLOps should care, as it provides a practical framework for effective AI-assisted research and highlights the critical human role in validating objectives.

7/10

Related reading

  1. The Agent Said It Was Done. The Database Disagreed.

    ThinkingBox benchmarks AI agents by checking the final database state after tool calls, revealing that many LLM‑driven agents succeed on a single attempt but fail to repeat the correct outcome. Across 507 tasks run 20 times, models differ widely in consistency and cost per successful attempt, with Kimi‑K3 being broad but inconsistent and Claude Opus models being more reliable.

    Hugging Facehuggingface.co13 minHN1
  2. Recursive self-improvement of AI research agents

    The paper introduces AIDE², an AI research agent that rewrites its own code, benchmarks each version, and adopts the best performing changes—a process they call recursive self‑improvement. In an 8‑day autonomous run it produced seven improvements that beat a strong human‑engineered baseline on four unseen benchmarks and reduced reward‑hacking from 55 % to 32 %.

    Hugging Face Daily Papersarxiv.org2 minpaperHN31
  3. Do LLMs Have the Memory of a Goldfish?

    The article explains that LLMs don’t have persistent personal memory; all “memory” is supplied by the surrounding application via the context window, summaries, or external storage. It outlines the distinction between trained weights, working‑memory (token context), and persistent application memory, shows how to construct API calls to preserve conversation state, and discusses the cost and laten…

    ByteByteGobytebytego.com12 min
  4. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

    Hugging Face Daily Papersarxiv.org2 minpaper
  5. EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks

    EMem‑Bench is a new benchmark of 2,554 long‑horizon embodied episodes that explicitly tests an agent’s ability to construct, update, and reuse memory across four defined challenges. The authors also release EMem, a spatial‑event‑scene external memory, and an 8B policy (EMem‑8B) that together achieve the strongest performance, highlighting persistent gaps in current multimodal LLM agents.

    Hugging Face Daily Papersarxiv.org1 minpaper