Copilot tops GitHub’s own AI code review benchmark. An independent one tells a different story.
The New Stack Paul Sawers
GitHub’s new AI code-review benchmark puts Copilot on top. An independent benchmark tells a messier story about who’s actually best.
Based on reporting by The New Stack, Paul Sawers — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
GitHub has added another scoreboard to the AI coding race. On Monday, it launched ReviewBench, an open benchmark built with Microsoft to test how well AI code-review agents spot useful problems in pull requests. The first winner is GitHub’s own Copilot code review, which sits atop the inaugural leaderboard in its Balanced mode.
ReviewBench uses 219 public pull requests from 187 public repositories across 19 programming languages. GitHub says it picked that set after analyzing 103.9 million pull requests, aiming for a mix that resembles GitHub overall rather than a pile of tiny, single-file changes. To define the right answers, it pulled from human review comments, later author changes, static-analysis tools and other LLM reviewers. Claude Sonnet 5 classifies the findings, and another LLM checks whether proposed issues match the reference set.
Copilot’s score is a 40.1% grounded F1, which blends precision and recall against the benchmark’s known findings. But the victory lap comes with a lot of footnotes. GitHub says its own team created the initial entries by running the public versions of the products, and the vendors did not conduct or verify those tests. The runs also happened on different dates, with Copilot tested on Oct. 1 and Cubic and Greptile back in June.
That matters because the benchmark may be pointing at a moving target. ReviewBench itself warns that products can change and that results on its corpus may not reflect a company’s own code. GitHub has still published the dataset, methodology and judging setup, and says vendors can submit their own runs. Alejandro Carderera de Diego, a staff applied engineer at GitHub, says the benchmark has already helped improve Copilot Code Review and has tracked the direction of later production experiments.
The release lands in a crowded field. Qodo, Greptile, Cubic, Devin, Cursor and CodeRabbit are all trying to automate more of code review, and Martian’s Code Review Bench already offers both online and offline leaderboards. On Martian’s online board, Cubic leads with 64.9% F1 and Copilot is fourth at 60.9%. On the offline board, Copilot is fifth with a 58% F2 score. The numbers aren’t directly comparable, but they do show the same thing: benchmark design can make the leaderboard look very different.
My take — AI-written commentary, not fact-checked reporting
A vendor ranking its own tool is the kind of benchmark news that should come with a tiny blinking warning light. GitHub may be serious about ReviewBench, but the real test is whether other people use it, not whether the home team gets to hoist the trophy first. Open data is nice; open competition is better.
Read more about this at: The New Stack
Related stories
Microsoft Open Sources code-testing-generator: a Polyglot Unit-Test Agent That Hits 92.1% Task Completion Versus 78.9% for Stock Copilot
MarkTechPost · 2 months ago ·
47
GitHub’s advice for its new Copilot feature is to try something else first
The New Stack · 5 days ago ·
5
AI hasn’t shifted the bottleneck from coding to code review
The New Stack · 2 months ago ·
32