The authors audit six popular physics benchmarks by having domain experts re‑grade model outputs, fixing reference answers and removing ambiguous items. After correction, GPT‑5.6‑Sol’s mean@4 jumps from ~47 % to ~79 % on HLE‑Physics and from ~61 % to ~87 % on CMT‑Benchmark, with a corrected pass@4 of 94 % on 54 vetted CritPt challenges. The work shows current benchmarks severely under‑report LLM…
Hacker News front page·arxiv.org·Ali Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He, Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'Hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. van den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo Cárdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. McIntosh, Evan McKinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker, Arman Cohan, John Sous··2 min·paperHN9650
Benchmark Radar is a continuously updated database and search engine that aggregates AI benchmark papers, datasets, and scores from 37 sources into a searchable catalog. It offers a web dashboard, CLI, and analysis of benchmark saturation to help LLM developers find and compare evaluations.
ImpossibleRubrics introduces a 169‑task benchmark of “impossible” prompts, each with a formal oracle certificate defining what an honest answer can claim. The authors generate rubrics downstream and test them adversarially, finding that many rubric generators are exploitable (8‑36% of the time) and that a single generic rubric (“be decisive, penalize hedging”) is exploited 64% of the time, while…
Android Bench 2.0 adds a set of long‑horizon tasks (multi‑day Android development problems) and introduces agent‑based evaluation. Scoring is now continuous, with the best model achieving a 28 % pass rate on these tasks, far lower than the ~91 % on earlier short tasks. The post lists new models on the leaderboard and points to updated methodology and GitHub repo.
Codename One engineers dissect why benchmark numbers can be misleading, then share concrete work on GC tuning, proper weak/soft references, and a new probing sequence for their open‑addressed HashMap that cuts miss‑probe counts from >16 k to ~1.5 per lookup.