TLDRocket
Sign in

Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors

MarkTechPost Asif Razzaq

Sakana AI built a review system that hunts for mistakes in papers, not just human-like comments. It found 73.43% of core-claim errors, but still got fooled by prompt injection.

Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sakana AI’s new TMLR paper, Beyond Imitation, takes a harder shot at AI peer review. Instead of asking whether a model can sound like a reviewer, the team asks whether it can spot a planted error in a paper. That sounds like a small shift. It isn’t. Once you change the target from imitation to detection, system design starts to matter a lot more than polished prose.

The setup is built around two pieces. One is a Contradiction Benchmark, which inserts contradictions into real papers. The other is Multi-Layered Review, or MLR, an agentic review system that reads a paper before judging it. MLR uses three agents running on off-the-shelf Claude models: one summarizes appendix details, one can do a literature check with web search, and one runs a three-pass review chain that produces strengths, weaknesses, questions, a recommendation, a score, and a to-do list. The PDF goes straight through, so figures and equations are kept in play.

The benchmark itself is unusually concrete. The researchers took 257 CC-licensed papers from ACL, AISTATS, CVPR, ICML 2025, and NeurIPS 2024. Gemini 2.5 Pro built a knowledge graph of claims, evidence, and methods. Then GPT-4.1 rewrote one node at each distance from the main claim into a contradiction, creating 1,164 inserted errors. An o3 judge scored the reviews ten times. On clean papers, it reached 99.9% accuracy, and it caught 86.8% of manually confirmed hits, which suggests the reported numbers may actually be cautious.

MLR led all four systems tested. With four reviews, it found 73.43% of core-claim errors and 40.95% overall on the benchmark, while the best baseline reached 14.81% on core claims. Even a single MLR review caught 60.79%. And this was not just a model swap story: the paper says changing GPT-4.1 to Claude Sonnet 4 inside a baseline lifted core-claim detection from 14.56% to 35.40%, while MLR’s design added about 25 more points on a single review.

The gains shrink on real retracted papers, which is the part that keeps the whole thing grounded. On WithdrarXiv-Check’s 211 papers, MLR scored 26.07% on similar matches and 16.11% on exact matches. It also correlated reasonably with human scores on ICLR 2025 submissions at 0.586, though humans agreed with each other more. The price is fairly low at about $0.47 per review, excluding the optional literature agent, and the authors say the system still falls for hidden prompt injection. That last detail matters more than the shiny headline.

My take — AI-written commentary, not fact-checked reporting

The useful lesson here is not that AI can replace reviewers; it’s that most “AI reviewer” demos were set up to flatter the model, not test it. A system that reads before judging is a better starting point than a chatbot with opinions, which is a low bar but apparently still a bar. The prompt-injection problem is the industry’s favorite reminder that clever evaluation can still get mugged by a sticky note in the input.

Read more about this at: MarkTechPost

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.