Hugging FaceAvijit Ghosh, Jenny Chim, Deep Joshi, Srishti, Matt Kennedy, Irene Solaiman, Jessica McFadyen, Lynn Tan, Coz3 min readintermediate
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Summary
The UK AI Security Institute (AISI) is publishing its frontier‑LLM benchmark results on EvalEval’s open Evaluation Cards platform, using the Every Eval Ever (EEE) schema. The release covers five main benchmarks (HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, Terminal‑Bench 2.0) and six models (Claude Opus 4/4.5/4.6, GPT‑5/5.2/5.4), plus two cyber‑evaluation suites. The data inclu…
- AISI’s evaluation results are now available as structured Evaluation Cards, providing raw run data, token‑budget curves, and benchmark metadata.
- The release demonstrates how inference compute and evaluation protocol affect performance (e.g., Humanity’s Last Exam token‑budget curves).
- EvalEval’s Every Eval Ever schema standardises benchmark reporting, making cross‑study comparisons possible.
- The collaboration showcases concrete tooling (EvalCards UI, EEE JSON schema) for reproducible LLM evaluation.
Reproducible evaluation is a bottleneck for LLM research; without standardized reporting, results can’t be reliably compared or audited. By publishing full run metadata, AISI provides reference points that expose how compute budgets and protocol choices bias scores, which is essential for both scie…
5/10


