Sakana AI’s LLM Peer Review System Catches 73% of Core-Claim Errors
MarkTechPost Asif Razzaq
Sakana AI built a review system that hunts for mistakes in papers, not just human-like comments. It found 73.43% of core-claim errors, but still got fooled by prompt injection.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI’s new TMLR paper, Beyond Imitation, takes a harder shot at AI peer review. Instead of asking whether a model can sound like a reviewer, the team asks whether it can spot a planted error in a paper. That sounds like a small shift. It isn’t. Once you change the target from imitation to detection, system design starts to matter a lot more than polished prose.
The setup is built around two pieces. One is a Contradiction Benchmark, which inserts contradictions into real papers. The other is Multi-Layered Review, or MLR, an agentic review system that reads a paper before judging it. MLR uses three agents running on off-the-shelf Claude models: one summarizes appendix details, one can do a literature check with web search, and one runs a three-pass review chain that produces strengths, weaknesses, questions, a recommendation, a score, and a to-do list. The PDF goes straight through, so figures and equations are kept in play.
The benchmark itself is unusually concrete. The researchers took 257 CC-licensed papers from ACL, AISTATS, CVPR, ICML 2025, and NeurIPS 2024. Gemini 2.5 Pro built a knowledge graph of claims, evidence, and methods. Then GPT-4.1 rewrote one node at each distance from the main claim into a contradiction, creating 1,164 inserted errors. An o3 judge scored the reviews ten times. On clean papers, it reached 99.9% accuracy, and it caught 86.8% of manually confirmed hits, which suggests the reported numbers may actually be cautious.
MLR led all four systems tested. With four reviews, it found 73.43% of core-claim errors and 40.95% overall on the benchmark, while the best baseline reached 14.81% on core claims. Even a single MLR review caught 60.79%. And this was not just a model swap story: the paper says changing GPT-4.1 to Claude Sonnet 4 inside a baseline lifted core-claim detection from 14.56% to 35.40%, while MLR’s design added about 25 more points on a single review.
The gains shrink on real retracted papers, which is the part that keeps the whole thing grounded. On WithdrarXiv-Check’s 211 papers, MLR scored 26.07% on similar matches and 16.11% on exact matches. It also correlated reasonably with human scores on ICLR 2025 submissions at 0.586, though humans agreed with each other more. The price is fairly low at about $0.47 per review, excluding the optional literature agent, and the authors say the system still falls for hidden prompt injection. That last detail matters more than the shiny headline.
My take — AI-written commentary, not fact-checked reporting
The useful lesson here is not that AI can replace reviewers; it’s that most “AI reviewer” demos were set up to flatter the model, not test it. A system that reads before judging is a better starting point than a chatbot with opinions, which is a low bar but apparently still a bar. The prompt-injection problem is the industry’s favorite reminder that clever evaluation can still get mugged by a sticky note in the input.
Read more about this at: MarkTechPost
Related stories
The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature
Sakana AI ·
6
An Introduction to AI Secure LLM Safety Leaderboard
Hugging Face · 2 years ago ·
13