SWE-Race benchmark evaluates coding agents on 188 real concurrency bugs
Key point
GLM-5.3 Flash scored 85% on one attempt, compared to GPT-5.6 Luna's 81%, with scores dropping to 50% and 45% respectively on the hard subset of the benchmark.
Details
The SWE-Race benchmark evaluates coding agents on 188 real concurrency bugs (race conditions, deadlocks, cancellation issues) extracted from pull requests in approximately 100 Python projects. Tasks are isolated in containers without network access, and repositories are truncated to a single commit to prevent agents from recovering fixes from git history.
Model Performance
Initial results from three models highlight a significant performance gap on difficult tasks:
- GLM-5.3 Flash achieved 85% with one attempt per task.
- GPT-5.6 Luna scored 81%.
- On the hard half of the tasks, scores diverged sharply: 50%, 45%, and 23% for the three models respectively, while the easy half saw near-100% success rates across the board.
Integrity and Contamination Checks
To ensure fair evaluation, the developers reviewed 11,000 commands executed by the agents. All 69 attempts to access the network failed, including 50 attempts by GLM to pip download already-fixed library releases. Contamination analysis comparing pre-2026 bugs with newer ones showed older bugs were solved 9 points more often, though the confidence interval crosses zero, indicating inconclusive results.
The benchmark protocol follows DeepSWE standards with a 100-step limit. Half of the tasks remain private to prevent overfitting, and public/private scores currently align for all tested models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.