TLDRocket
Sign in

Can AI automate computational reproducibility?

AI as Normal Technology Sayash Kapoor

Researchers introduced CORE-Bench, a benchmark for measuring how well AI agents can automate computational reproducibility in scientific research. The best AI agent tested (CORE-Agent with GPT-4o) achieved 22% accuracy on the hardest difficulty level despite task-specific modifications. The work suggests that AI systems may prove economically valuable for automating specific scientific tasks even without general-purpose capabilities, challenging conventional notions of artificial general intelligence.

Why it matters

A new benchmark to measure the impact of AI on improving science

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.