TLDRocket
Sign in

Agent Benchmarking

21 summarised stories about Agent Benchmarking, each linking back to the original source. Browse all topics →

Thursday, 30 April 2026

AstaBench update: New results, plus adoption from industry

Allen Institute (AI2) 2 months ago

AstaBench, an open benchmark for measuring AI agent scientific research capabilities, released new results testing frontier models including GPT-5.5 on over 2,400 research problems and updated its leaderboard. Claude Opus 4.7 achieved the top overall score of 58.0% while GPT-5.5 reached 52.9% at $1.61 per problem, showing that frontier models improve unevenly across categories with Code & Execution and End-to-End Discovery gaining substantially but Data Analysis and Literature Understanding gaining only moderately. The benchmark is gaining adoption from the UK AI Security Institute, General Reasoning, and other organizations, establishing AstaBench as an industry standard for evaluating AI's ability to perform grounded scientific research workflows.

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.