AivexaNewsSearch
AI news for builders and product teamsChecked every hour

ReviewBench: An open benchmark for AI code review

Collected Oct 5, 2026

GitHub has introduced ReviewBench, an open offline benchmark for AI code review, and it is available for use today. The company says it built the benchmark to address a gap in existing evaluation methods, which often trade off label quality, coverage, and real-world representation.

ReviewBench follows the language, repository size, and size distribution of pull requests, modeled after more than 100 million real pull requests on GitHub. It uses a multi-source golden set—a validated collection of known findings for each pull request—and a consistent evaluation rubric, and it has been independently validated by senior engineers. The benchmark supports breakdowns by severity, category, and precision-recall preferences, and provides an offline signal intended to track whether changes are likely to improve production experiences.

GitHub reports that its offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments with the help of ReviewBench. The company also states that the benchmark is designed for teams building code review agents and that they can onboard their own systems and submit results.

Why it matters: Developers and product teams working on AI code review can use ReviewBench to compare what different systems catch, what they miss, and the tradeoffs they make, using a reproducible methodology.

Read at GitHub Blog · AI & ML

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We’re launching ReviewBench, a benchmark for code review agents built on representative GitHub pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog .