Expert Re-grading Confirms Frontier Model Performance Underestimated Due to Physics Benchmark Flaws
Key point
Expert re-grading confirmed that frontier models' actual capabilities were underestimated due to scoring errors and flaws in existing physics benchmarks.
Details
It has been revealed that the low scores recorded by frontier models on existing physics benchmarks were due to flaws in the evaluation system itself, rather than a lack of model capability. A team of approximately 50 researchers, including Ali Ansari, conducted an expert audit of six major physics benchmarks, identifying and correcting scoring errors, incorrect reference answers, and ambiguous problems.
After correcting errors and excluding flawed problems, re-evaluation results showed a significant increase in model performance. In particular, GPT-5.6-Sol's mean@4 score on HLE-Physics surged from 47.3% to 78.7%, and on CMT-Benchmark it rose from 61.0% to 87.2%. On the 54 retained problems in CMT-Benchmark, pass@4 reached 94.4%.
These results suggest that existing benchmarks significantly underestimate frontier models' capabilities on well-defined physics problems. The researchers noted that model performance on closed-form tasks has reached near saturation and emphasized the need for more rigorous, expert-validated evaluation standards.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.