AI Briefing
KO

SWE-bench Verified No Longer Measures Frontier Coding Capability

·2026.04.07 18:06

Key point

OpenAI pointed out contamination and test flaws in SWE-bench Verified and recommended discontinuing it.

Details

OpenAI stated that SWE-bench Verified no longer properly reflects the actual capabilities of frontier code models. Amid slowing performance gains, they explained this conclusion came from re-examining whether remaining failures were model limitations or dataset problems.

There are two main issues.

  • Test flaws: Among the 138 problems audited, tests were incorrect in 59.4%, meaning even functionally correct solutions could fail.
  • Training contamination: Frontier models likely encountered the problems or correct patches during training, and in fact some models reproduced the gold patch or details of the problem description verbatim.

OpenAI presented two types of failure as examples.

  • narrow test cases: Force specific implementation details, over-constraining problems that could have multiple valid solutions
  • wide test cases: Check for additional functionality not mentioned in the problem description, disqualifying models that fixed the issue exactly as described

To address the limitations of the initial SWE-bench, OpenAI created the SWE-bench Verified 500-problem set in 2024 through expert review, but this analysis concluded that even that version is unsuitable for evaluating current-level frontier models.

OpenAI has now discontinued reporting SWE-bench Verified scores, and stated it will create a new, uncontaminated evaluation instead. In the meantime, they recommended reporting SWE-bench Pro results.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.