AstaBench update: New results, plus adoption from industry
Allen Institute (AI2)
AstaBench, an open benchmark for measuring AI agent scientific research capabilities, released new results testing frontier models including GPT-5.5 on over 2,400 research problems and updated its leaderboard. Claude Opus 4.7 achieved the top overall score of 58.0% while GPT-5.5 reached 52.9% at $1.61 per problem, showing that frontier models improve unevenly across categories with Code & Execution and End-to-End Discovery gaining substantially but Data Analysis and Literature Understanding gaining only moderately. The benchmark is gaining adoption from the UK AI Security Institute, General Reasoning, and other organizations, establishing AstaBench as an industry standard for evaluating AI's ability to perform grounded scientific research workflows.
Why it matters
AstaBench’s latest update adds new frontier-model results, including GPT-5.5, and highlights growing adoption from groups including the UK AISI, General Reasoning, Elicit, SciSpace, Distyl AI, and EvoScientist.