Evaluating agents for scientific discovery
Allen Institute (AI2)
AI2 released two benchmarks—ScienceWorld and DiscoveryWorld—to evaluate whether AI agents can perform scientific tasks rather than just answer questions about science. In DiscoveryWorld's 120 challenge tasks, top models complete only around 20% at higher difficulty levels while human scientists with advanced degrees solve approximately 70%. The benchmarks reveal a gap between knowing scientific concepts and applying them through experimentation, with the tools now freely available as AI agent development accelerates.
Why it matters
Two benchmarks developed at Ai2 – ScienceWorld and DiscoveryWorld – reveal that even incredibly strong AI science agents struggle with problems human scientists solve routinely.