AI Briefing
KOSign in

Researchers Release SWE-sweep Benchmark for Proactive Bug Detection in Codebases

·2026.10.03 01:05

Key point

The new MIT-licensed benchmark evaluates models on finding and fixing real-world bugs in large codebases before users encounter them.

Details

Researchers from Meta, Stanford, Harvard, and UW have released SWE-sweep, a new benchmark designed to test whether language models can proactively find and fix bugs in large codebases without prior user reports.

Benchmark Methodology

Unlike traditional benchmarks that ask models to fix a specific bug already encountered by a user, SWE-sweep provides an agent with an entire codebase and asks it to identify and resolve as many issues as possible. The scoring is based on a hidden set of known real-world bugs within the repositories. The team applied rigorous filtering to ensure all included bugs are discoverable and fixable solely by reading the code.

Initial Results and Availability

Early results indicate that Luna xhigh currently offers the best cost-efficiency, though overall scores are lower than anticipated. While some tasks, such as fixing issues in numpy, are performed at a superhuman level, performance drops significantly on smaller repositories. The project is fully open-source under the MIT license, with code and paper available on GitHub and the project website.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.