AI Briefing
KOSign in

GitHub releases ReviewBench, an open benchmark for AI code review agents

·2026.10.06 00:59

Key point

GitHub has released ReviewBench, an open offline benchmark for AI code review agents, constructed from an analysis of 103.9 million GitHub pull requests.

Details

GitHub has released ReviewBench, an open offline benchmark designed to evaluate AI code review agents. The tool addresses the difficulty of measuring reviewer quality by providing a standardized evaluation framework that reflects real-world pull request distributions. It is available immediately for teams to assess their own agents and compare performance against others.

Benchmark Construction and Scope

ReviewBench was built by analyzing the language, repository size, and change shape of 103.9 million real GitHub pull requests. The resulting benchmark corpus consists of 219 public pull requests from 187 open-source repositories spanning 19 languages. While the language and repository-size distributions mirror GitHub overall, the pull request size distribution is weighted toward substantive, multi-file changes to avoid overrepresenting tiny, single-file edits.

The benchmark uses a multi-source golden set derived from human reviewers, author follow-up commits, static analysis tools, and multiple frontier LLMs. These findings are semantically deduplicated and validated under a shared rubric. Senior engineers independently re-labeled the ground-truth findings before release, achieving a 96.6% agreement rate with the benchmark's labels.

Evaluation Metrics and Flexibility

ReviewBench reports six metrics divided into two families to capture both known and newly discovered issues:

  • Grounded metrics (precision, recall, F1): Compare agent findings strictly against the existing golden set.
  • Augmented metrics (precision, recall, F1): Allow the judge to validate findings not present in the golden set, giving credit for valid issues that the original creators missed.

The benchmark supports configurable evaluation, allowing users to slice results by severity (Critical, Medium, Low) and category (Correctness, Security, Reliability, etc.). Users can also adjust the Fβ score to prioritize either precision (less noise) or recall (broader coverage), with the leaderboard re-ranking accordingly.

Impact on Copilot Code Review

GitHub used ReviewBench to evaluate successive iterations of Copilot code review (CCR). The benchmark provided an offline signal that reliably predicted the direction of production A/B tests. In a recent lite-tier experiment involving a multi-model ensemble review, ReviewBench predicted improvements in precision, recall, and comment volume, along with lower costs.

The subsequent online A/B test confirmed these predictions:

  • Addressed rate (online precision proxy) rose 8.0%.
  • Recall rose 13.6%.
  • Comment volume rose 61%.
  • Cost per review fell 8.0%.

ReviewBench also accurately predicted a 227% increase in critical comments, closely matching the 262% increase observed online. This demonstrates the benchmark's utility in providing a fast, repeatable signal before committing to production experiments.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.