Can AI automate computational reproducibility?
AI Snake Oil Sayash Kapoor
Researchers introduced CORE-Bench, a benchmark for measuring how well AI agents can automate computational reproducibility in scientific research. The best AI agent tested (CORE-Agent with GPT-4o) achieved 22% accuracy on the hardest difficulty level despite task-specific modifications. The work suggests that AI systems may prove economically valuable for automating specific scientific tasks even without general-purpose capabilities, challenging conventional notions of artificial general intelligence.
Why it matters
A new benchmark to measure the impact of AI on improving science