proomt

Search

Search posts, papers, and topics

All posts

ByteByteGo1 min readrelease notesintro

Last 3 days: AI Evals, October cohort

Summary

This post announces a live, hands-on course on building reliable evaluation systems for production AI agents, with enrollment closing soon. The course covers designing evals for quality, safety, cost, and latency, and building LLM-as-a-Judge systems.

  • Learn to design evals for quality, safety, reliability, cost, and latency.
  • Build and validate LLM-as-a-Judge systems for AI evaluation.
  • Red-team agents for prompt injection and jailbreaks.
  • Create meaningful eval datasets from real and synthetic data.
2/10

Related reading

  1. Agentic Skill Decay

    Addy Osmani warns that AI agents can short‑circuit the hands‑on practice (“reps”) that builds deep expertise and judgment. He recommends deliberately inserting hypothesis‑forming, “why” questioning, diff inspection, failure prediction, and occasional manual coding into the workflow, especially for junior engineers. A 2026 Anthropic study showed junior developers using AI scored 17 % lower on a fo…

    Addy Osmaniaddyosmani.com16 min
  2. AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop?

    AutoDataBench introduces a new benchmark to evaluate if AI agents can autonomously generate high-quality training data for LLMs, judged against production-like acceptance criteria. The study found that current agents can produce usable tasks but struggle significantly with efficiency, scoring low on a time-constrained budget.

    Hugging Face Daily Papersarxiv.org2 minpaper
  3. How UK AISI and EvalEval Are Making Benchmark Results Reproducible

    The UK AI Security Institute (AISI) is publishing its frontier‑LLM benchmark results on EvalEval’s open Evaluation Cards platform, using the Every Eval Ever (EEE) schema. The release covers five main benchmarks (HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, Terminal‑Bench 2.0) and six models (Claude Opus 4/4.5/4.6, GPT‑5/5.2/5.4), plus two cyber‑evaluation suites. The data inclu…

    Hugging Facehuggingface.co3 min
  4. Presentation: Teaching Engineers, Trusting AI: How Education Enabled Autonomous Code Review

    Duolingo’s DevEx AI team built a program of AI‑literacy workshops, observability dashboards, office‑hours, and vendor partnerships to get engineers comfortable with LLM‑based tools. With that foundation they launched a PR‑risk‑assessment bot that auto‑approves low‑risk pull requests, cutting review bottlenecks while keeping defect rates flat.

    InfoQinfoq.com24 mintalk
  5. Prompts aren’t Real

    The talk argues that prompt engineering is a dead‑end and proposes building large evaluation/optimization pipelines (pass^k testing, adversarial scenario generation, automated prompt optimization) to make LLM agents reliable. It describes a workflow: generate tests, run them with/without a new “skill”, feed results to a genetic optimizer that mutates prompts, validate on hold‑out tests, and itera…

    Hacker News front pageevaluation.club24 mintalkHN11757
  6. AutoSynthData: Generating Training Data for Enterprise Agents

    AutoSynthData generates synthetic training data for enterprise AI agents by identifying a target model's weaknesses and using a stronger teacher to guide task creation. It produces feasible, realistic, and difficult tasks, validated through execution and verification, to iteratively improve agent performance.

    Hugging Facehuggingface.co9 min