Evaluating agents for scientific discovery
Allen Institute (AI2)
AI2's benchmarks show AI 'science agents' still flunk most hands-on experiments, even though they ace science trivia. Everyone's hyping science agents, but the actual test scores tell a much less impressive story.
Based on reporting by Allen Institute (AI2) — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
The gap between an AI that can quote the boiling point of water and one that can actually figure it out with a thermometer and a stove turns out to be enormous. That's the blunt lesson from two benchmarks built at the Allen Institute for AI, ScienceWorld and DiscoveryWorld, which are suddenly getting a lot more attention now that frontier models are finally strong enough to make a dent in them.
ScienceWorld, released in 2022, drops an agent into a simulated world with a kitchen, a greenhouse, a workshop, and roughly 200 objects that behave like the real thing: ice melts, circuits conduct, plants grow if you treat them right. The twist is that models which scored top marks on multiple-choice grade-school science exams back then scored below 10% on ScienceWorld's 30 task types. Three years later, per Microsoft Research's 2025 TALES benchmark suite, leading models have climbed into the low 80s. That is genuine progress. But it still falls short of a full 4th-grade curriculum, which is a strange thing to type about systems that can write essays on quantum field theory.
DiscoveryWorld, added in 2024, raises the stakes considerably. Set on a fictional space colony called Planet X, it hands agents 120 tasks across eight domains, from proteomics to radioisotope dating, and asks them to form hypotheses, design experiments, run them, and interpret the results from scratch. Everything is fictional on purpose, so an agent can't cheat by pattern-matching to something it memorized during training. Human scientists with advanced degrees solve the hardest tasks about 70% of the time. The best current agents manage roughly 20%.
Ai2 researcher Peter Jansen, who built both benchmarks, put it plainly: if the top systems a year ago couldn't clear the easy tier, there's no strong reason to assume this year's crop of demo videos represents some quiet leap forward. His point isn't that agents are useless. It's that the loud claims flooding social feeds, about agents writing full papers or independently designing experiments, are rarely backed by anything resembling this kind of scored, repeatable test. ScienceWorld followed the same arc when it launched, starting brutally hard before models eventually caught up. Jansen thinks DiscoveryWorld is entering that phase now, which is either an encouraging sign of momentum or a reminder that today's headlines are running well ahead of the underlying capability.
Both benchmarks are free and open, which matters more than it sounds. If AI is eventually going to help cure diseases or discover new materials, as Jansen hopes, the only way to know it's ready is to keep handing it problems it can't talk its way around.
My take — AI-written commentary, not fact-checked reporting
I'll believe the science-agent hype the day one of these systems clears DiscoveryWorld's hard tier without a human quietly nudging the setup, because right now the scoreboard says 20% versus a PhD's 70%. The pattern here is depressingly familiar from the LLM era generally: impressive demos, thin evaluation, and a benchmark that quietly does the unglamorous work of keeping everyone honest. More credit to AI2 for building the boring, rigorous thing instead of the flashy one.
Read more about this at: Allen Institute (AI2)