AI Briefing
KO

Errors Found in 12% of Major AI Benchmarks, Cleaned Versions Released

·2026.07.29 04:58

Key point

About 12% of questions in major AI benchmarks such as GPQA and MMLU-Pro were found to contain errors, leading to the release of cleaned datasets.

Details

A full review of the GPQA (Diamond/Extended), MMLU-Pro, and MMMU-Pro benchmarks found that about 12% of the questions had defects, including question construction errors, incorrect answer keys, and cases with multiple possible correct answers.

Along with corrections for these errors, a Clean version dataset and a detailed list (ledger) showing which questions were defective and why have been released together. When testing was conducted after removing the defective questions, the performance of top models rose from the existing 92-93% level to about 98%.

The released resources are as follows:

  • Cleaned benchmark datasets (Hugging Face)
  • Tasks for lm-eval-harness
  • Flagged-candidate ledger recording the causes of errors
  • A paper and GitHub repository containing detailed analysis

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.