TLDRocket
Sign in

AstaBench update: New results, plus adoption from industry

Allen Institute (AI2)

AstaBench, an open benchmark for measuring AI agent scientific research capabilities, released new results testing frontier models including GPT-5.5 on over 2,400 research problems and updated its leaderboard. Claude Opus 4.7 achieved the top overall score of 58.0% while GPT-5.5 reached 52.9% at $1.61 per problem, showing that frontier models improve unevenly across categories with Code & Execution and End-to-End Discovery gaining substantially but Data Analysis and Literature Understanding gaining only moderately. The benchmark is gaining adoption from the UK AI Security Institute, General Reasoning, and other organizations, establishing AstaBench as an industry standard for evaluating AI's ability to perform grounded scientific research workflows.

Why it matters

AstaBench’s latest update adds new frontier-model results, including GPT-5.5, and highlights growing adoption from groups including the UK AISI, General Reasoning, Elicit, SciSpace, Distyl AI, and EvoScientist.

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.