The AI Scientist Generates its First Peer-Reviewed Scientific Publication
Sakana AI ● Covered by 2 sources
Sakana AI's automated system wrote a full research paper that passed peer review at an ICLR workshop, with zero human edits. Reviewers didn't know which papers were AI-made, and this one scored better than most human submissions.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI just crossed a line that a lot of researchers assumed was still years away. The AI Scientist-v2, the company's autonomous research system, produced a paper end to end — hypothesis, experiments, code, analysis, figures, every sentence of the writeup — and submitted it to an ICLR 2025 workshop under a double-blind review arrangement the company set up with the workshop organizers and ICLR leadership, plus IRB approval from the University of British Columbia.
Three AI-generated papers went into the pile of 43 submissions. Reviewers were told some papers might be machine-written, but not which ones. Two of the three flopped. The third, titled "Compositional Regularization: Unexpected Obstacles in Enhancing Neural Network Generalization," pulled an average score of 6.33, which put it around the 45th percentile of all submissions and above the workshop's acceptance line. It reported a negative result — the system tried to improve how neural networks generalize compositionally and largely failed, then wrote that failure up as a paper.
Here's the twist: Sakana pulled it before publication anyway. That decision was baked into the experiment's protocol from the start, not a reaction to the outcome. The team and the workshop agreed the paper would be withdrawn and desk-rejected regardless of the score, because nobody has settled the norms yet for whether AI-authored manuscripts belong in the same venues as human work. No public OpenReview listing, no proceedings entry.
Sakana's own researchers also ran a parallel, harsher review, treating all three papers as if they'd been submitted to ICLR's main conference track rather than the workshop. None passed that bar. They found real flaws too — including a citation blunder where the system credited an LSTM description to Goodfellow's 2016 deep learning book instead of Hochreiter and Schmidhuber's original 1997 paper. Workshop acceptance rates at venues like this typically run 60-70%, versus 20-30% at the main conference, so clearing the workshop bar is a meaningfully lower hurdle than clearing ICLR itself. Sakana is upfront that the system's ceiling is tied to how good the underlying large language models get, and expects both to keep climbing.
My take — AI-written commentary, not fact-checked reporting
Getting past a workshop's review bar is not the same as getting into the main conference, and Sakana deserves credit for saying so plainly instead of spinning a 6.33 average score into more than it is. Pulling the paper before publication was the right call, but it also means the actual test — would the wider scientific community accept AI-authored work into its permanent record — never happened. That question is still sitting there unanswered, and it is the one that actually matters.
Read more about this at: Sakana AI
Related stories
The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature
Sakana AI ·
3
Improving the academic workflow: Introducing two AI agents for better figures and peer review
Google Research · 5 months ago ·
2
PaperBench: Evaluating AI’s Ability to Replicate AI Research
OpenAI · 1 year ago ·
3