The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature
Sakana AI ● Covered by 4 sources
Sakana AI's paper on its AI Scientist system just got published in Nature, after the tool already passed real peer review last year. An automated reviewer they built now judges AI papers about as well as humans do.
Based on reporting by Sakana AI — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Sakana AI has spent roughly a year and a half building something it calls The AI Scientist, and this week that work landed in Nature, backed by researchers from the University of British Columbia, the Vector Institute, and Oxford. The pitch is blunt: an agent that comes up with a research idea, reads the relevant literature, designs and runs experiments through a parallelized search process, and then writes the whole paper in LaTeX, complete with a vision-capable model checking its own figures.
The headline result isn't new exactly, but the paper puts real numbers behind it. An earlier version, AI Scientist-v2, already produced a fully AI-generated paper that passed peer review at a workshop for a top-tier AI conference. What's new here is the machinery Sakana built to check that kind of claim without burning out human reviewers: an Automated Reviewer that acts like an Area Chair, combining five independent reviews into one verdict based on NeurIPS guidelines. Tested against thousands of real decisions from OpenReview, it hit a balanced accuracy of 69%, on par with human reviewers, and its F1-score actually beat the inter-human agreement measured in the well-known NeurIPS 2021 consistency study.
The more interesting finding is what happens when you swap in stronger foundation models underneath. Using the Automated Reviewer as a consistent judge, Sakana found that paper quality scales cleanly with the capability of the underlying model. Better foundation models produce better papers, full stop. That's the kind of scaling relationship that tends to keep paying off as compute gets cheaper and models keep improving, which is presumably why Sakana is comfortable calling this early days rather than a finished product.
And it is early. The paper is candid about current limitations, and the system is still confined to computational experiments rather than, say, wet-lab science. But Sakana's argument is that machine learning capabilities have a habit of going from barely-working to superhuman fast once the core loop is proven out, and they expect the same pattern to eventually spread beyond computation.
The ethical questions get real estate in the paper too. Automating paper generation raises the risk of flooding peer review systems or padding research credentials that were never meant to be automated. Sakana says it withdrew its own accepted AI-written submissions, got IRB approval before running these experiments on real review systems, and now watermarks every AI-generated paper it puts out, a practice it's pushing the rest of the field to adopt as well.
My take — AI-written commentary, not fact-checked reporting
The Nature stamp gives this story more institutional weight than a preprint ever could, but the genuinely useful contribution here is the Automated Reviewer, not the paper-writing agent itself. Anyone can hype up an AI that writes convincing papers; building a review judge that outperforms measured human agreement on the same task is a much harder thing to fake. The watermarking and IRB moves are the right instinct, and the rest of the field should copy them before someone tries to sneak an unmarked AI paper past a conference that isn't looking for one.”
Read more about this at: Sakana AI
Related stories
The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
Sakana AI ·
42
The AI Scientist Generates its First Peer-Reviewed Scientific Publication
Sakana AI ·
20
Improving the academic workflow: Introducing two AI agents for better figures and peer review
Google Research · 5 months ago ·
2