Introducing SWE-bench Verified
Key point
OpenAI has released **SWE-bench Verified**, a human-validated benchmark, to more accurately evaluate AI's software engineering capabilities.
Details
As part of the Preparedness Framework, which tracks models' autonomous behavior capabilities, OpenAI has released SWE-bench Verified, which enables more precise evaluation of AI's software engineering capabilities. This is a human-validated dataset created to address the issue of the existing SWE-bench underestimating models' actual capabilities.
The existing SWE-bench measures AI agent performance using GitHub issues and PRs, but three major limitations were found:
- Unit tests are overly specific or unrelated to the issue, causing correct solutions to be marked as incorrect
- Issue descriptions are insufficient, making the intent of the problem ambiguous
- Development environment setup issues cause valid code to be marked as failing
SWE-bench Verified improves upon these issues, helping to more reliably measure AI agents' ability to solve real-world software problems.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.