GitHub’s ReviewBench: A Shared Scorecard for AI Code Reviewers

Published:

More and more software is now checked by AI before it ships. Tools that read a programmer’s proposed changes and flag problems are becoming routine. But until now there has been no easy, open way to tell whether one AI reviewer is actually better than another.

GitHub wants to change that. On 5 October 2026 it published ReviewBench, an open benchmark for AI code-review agents built from real pull requests, and invited anyone to test their own tool against it [1].

What is ReviewBench?

A benchmark is a standard test: every system gets the same set of tasks and is marked the same way, so results can be compared fairly.

ReviewBench’s tasks are pull requests, the bundles of proposed code changes that developers submit for review before they are merged into a project. GitHub says it analysed 103.9 million pull requests on its platform to understand what typical review work looks like, then built a test set of 219 public pull requests from 187 open-source repositories, covering 19 programming languages [1].

GitHub says the language and repository-size mix closely matches GitHub overall, but it deliberately tilted the set away from tiny, single-file changes towards larger ones where review quality matters most [1]. The full dataset is public.

How are reviewers marked?

To mark answers, you need an answer key. GitHub calls this the “golden set”: a checked list of genuine problems in each pull request.

The company says it gathered candidate findings from several sources, including human reviewers, frontier AI models and static-analysis tools (programs that scan code for known kinds of mistakes) [1]. Duplicates were merged, and each finding was checked against one shared rubric. An AI model, Claude Sonnet 5, acts as the grader, and GitHub says it publishes both the rubric and the grader set-up [1].

Results are reported in two ways:

  • Grounded scores count only matches against the known answer key.
  • Augmented scores also give credit when a reviewer finds a real problem the answer key missed, with the grader deciding whether it is genuine.

Both use precision (how much of what a tool flags is actually valid, so less noise) and recall (how many of the known problems it finds). GitHub uses grounded recall as its headline comparison [1]. Users can also re-weight the leaderboard depending on whether they care more about catching everything or about avoiding false alarms.

How reliable is it?

GitHub says that before release it asked senior engineers who had not helped build the dataset to re-label every answer-key finding from scratch, and their judgements agreed with ReviewBench 96.6% of the time [1]. That figure is GitHub’s own audit, not an outside review.

The company also says ReviewBench has helped it develop Copilot code review: in its experience, changes that score better offline have consistently pointed the same way as later live tests with real users [1]. It gives one example, an experiment combining several model runs into a single review, where it says the live results moved in the direction the benchmark predicted.

Teams can submit their own agents, but scores stay private until a maintainer reviews and approves them for the public leaderboard [1].

What we don’t know

  • Grader bias. Using one AI model as the marker, and building the answer key partly from AI output, could tilt scores. GitHub acknowledges this and publishes its known threats to validity [1].
  • Independent checking. The 96.6% agreement figure and the offline-versus-live track record are GitHub’s own results.
  • Who controls the leaderboard. Publication depends on maintainer approval, and GitHub runs a competing product.

The Bottom Line

ReviewBench gives AI code review something it has lacked: a public, reproducible test built from real-world work. That is useful for anyone choosing or building a review tool. But it was built and marked by GitHub, using an AI grader, so outside researchers poking at its rubric and results will decide how much weight it deserves.

Sources

  1. GitHub Blog, Michelle Zhou and Alejandro Carderera de Diego, 5 October 2026: announcement of ReviewBench, an open benchmark for AI code review. https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/
TSN
TSNhttps://tsnmedia.org/
Welcome to TSN. I'm a data analyst who spent two decades mastering traditional analytics—then went all-in on AI. Here you'll find practical implementation guides, career transition advice, and the news that actually matters for deploying AI in enterprise. No hype. Just what works.

Related articles

Recent articles