AI Briefing
KO
Pick

SWE-Bench Pro V2 Released: 89 Tasks Removed and Evaluation Protocol Strengthened

·2026.09.24 02:15

Key point

SWE-Bench Pro V2, co-developed with Reflection, has been released, restructured to 642 tasks after removing 89 problematic ones identified during verification, with an enhanced evaluation protocol.

Details

SWE-Bench Pro V2 is a benchmark for evaluating the software engineering performance of AI agents, co-developed with Reflection. In this version, 89 tasks identified as problematic during verification were removed, restructuring the benchmark to a total of 642 tasks. Regarding the strengthened evaluation protocol, access to web tools and external repositories was blocked to prevent models from searching for correct answers, and agent access to model endpoints was restricted. Additionally, all patches are rigorously verified against reference implementations, with improved dependency support applied to 211 tasks and validator error fixes applied to 10 tasks. Key changes mentioned include the handling of 69 tasks where instructions and tests were contradictory, and remaining issues where code within patches (such as conftest.py) is still executed by the validator. According to the latest leaderboard, Opus 5 (Claude Code) ranks first with a score of 98.00, Fable 5.1 (Claude Code) ranks second with 92.20, and GPT-6 (Codex) ranks third with 90.20.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.