Hugging Face Daily PapersKoutian Wu, Junjie Zhou, Ergan Shang1 min readpaperintermediate
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
Summary
Benchmark Radar is a continuously updated database and search engine that aggregates AI benchmark papers, datasets, and scores from 37 sources into a searchable catalog. It offers a web dashboard, CLI, and analysis of benchmark saturation to help LLM developers find and compare evaluations.
- Aggregates daily from 37 sources (13 connectors, 24 feeds) into a searchable benchmark catalog.
- Catalog holds 1,283 source records and 12,916 numeric observations across 790 benchmarks.
- Web UI includes leaderboard, Pareto frontier, saturation/trend views; a CLI enables offline queries.
- Audit reveals benchmark saturation and adoption trends, exposing limits of direct score comparisons.
LLM researchers and evaluation engineers should care because it centralizes benchmark data and tools for systematic discovery and comparison.
7/10


