TLDRocket
Sign in

The AI Scientist: Towards Fully Automated AI Research, Now Published in Nature

Sakana AI Covered by 4 sources

Sakana AI's paper on its AI Scientist system just got published in Nature, after the tool already passed real peer review last year. An automated reviewer they built now judges AI papers about as well as humans do.

Based on reporting by Sakana AI — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sakana AI has spent roughly a year and a half building something it calls The AI Scientist, and this week that work landed in Nature, backed by researchers from the University of British Columbia, the Vector Institute, and Oxford. The pitch is blunt: an agent that comes up with a research idea, reads the relevant literature, designs and runs experiments through a parallelized search process, and then writes the whole paper in LaTeX, complete with a vision-capable model checking its own figures.

The headline result isn't new exactly, but the paper puts real numbers behind it. An earlier version, AI Scientist-v2, already produced a fully AI-generated paper that passed peer review at a workshop for a top-tier AI conference. What's new here is the machinery Sakana built to check that kind of claim without burning out human reviewers: an Automated Reviewer that acts like an Area Chair, combining five independent reviews into one verdict based on NeurIPS guidelines. Tested against thousands of real decisions from OpenReview, it hit a balanced accuracy of 69%, on par with human reviewers, and its F1-score actually beat the inter-human agreement measured in the well-known NeurIPS 2021 consistency study.

The more interesting finding is what happens when you swap in stronger foundation models underneath. Using the Automated Reviewer as a consistent judge, Sakana found that paper quality scales cleanly with the capability of the underlying model. Better foundation models produce better papers, full stop. That's the kind of scaling relationship that tends to keep paying off as compute gets cheaper and models keep improving, which is presumably why Sakana is comfortable calling this early days rather than a finished product.

And it is early. The paper is candid about current limitations, and the system is still confined to computational experiments rather than, say, wet-lab science. But Sakana's argument is that machine learning capabilities have a habit of going from barely-working to superhuman fast once the core loop is proven out, and they expect the same pattern to eventually spread beyond computation.

The ethical questions get real estate in the paper too. Automating paper generation raises the risk of flooding peer review systems or padding research credentials that were never meant to be automated. Sakana says it withdrew its own accepted AI-written submissions, got IRB approval before running these experiments on real review systems, and now watermarks every AI-generated paper it puts out, a practice it's pushing the rest of the field to adopt as well.

My take — AI-written commentary, not fact-checked reporting

The Nature stamp gives this story more institutional weight than a preprint ever could, but the genuinely useful contribution here is the Automated Reviewer, not the paper-writing agent itself. Anyone can hype up an AI that writes convincing papers; building a review judge that outperforms measured human agreement on the same task is a much harder thing to fake. The watermarking and IRB moves are the right instinct, and the rest of the field should copy them before someone tries to sneak an unmarked AI paper past a conference that isn't looking for one.”

Read more about this at: Sakana AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.