Hacker News front pageAli Ansari, Haoran Sun, Andy Zeyi Liu, Mark Jabbour, Yongshan Ding, Steven Girvin, Yu He, Sohrab Ismail-Beigi, Aleksander Kubica, Owen D. Miller, Corey O'Hern, Vidvuds Ozolins, David Poland, A. Douglas Stone, Frank C. van den Bosch, Logan Wright, Navid Akbari, Santanu Antu, Kangle Cai, Andrew Calabrese-Day, Mateo Cárdenes Wuttig, Meng Cheng, Barry T. Chiang, Ali Ghorashi, Shouzhen Gu, Haoyang Huang, Zhibo Kang, Lukas Kienesberger, Hantian Liu, Charles Lomba, Zhongling Lu, Wenchao Ma, Rohin E. McIntosh, Evan McKinney, Ivan Rojkov, Xulei Sun, Yarone Meir Tokayer, Naveen Balaji Umasankar, Mira Varma, Leda Wang, Qimin Wang, Tyler Wang, Haoyu Wei, Jinming Yang, Jinchen Zhao, Sherlock Tingrui Zhao, Qinyuan Zheng, Jay S. Zou, Lucas Baker, Arman Cohan, John Sous2 min readpaperadvanced
How good are frontier models at physics?
Summary
The authors audit six popular physics benchmarks by having domain experts re‑grade model outputs, fixing reference answers and removing ambiguous items. After correction, GPT‑5.6‑Sol’s mean@4 jumps from ~47 % to ~79 % on HLE‑Physics and from ~61 % to ~87 % on CMT‑Benchmark, with a corrected pass@4 of 94 % on 54 vetted CritPt challenges. The work shows current benchmarks severely under‑report LLM…
- Many “incorrect” model answers were actually due to flawed benchmark design—wrong reference solutions, ambiguous questions, or grading errors.
- Expert re‑grading can raise reported scores dramatically (e.g., GPT‑5.6‑Sol mean@4 from 47.3 % → 78.7 % on HLE‑Physics).
- Closed‑ended physics tasks are nearing saturation for frontier models, indicating the need for more challenging, well‑curated evaluation sets.
If benchmark scores are biased low, research and product decisions based on them (e.g., model selection, claims of progress) may be misguided. Reliable, expert‑validated evaluations are essential for measuring true scientific reasoning in LLMs and for guiding future model development.
6/10
