TLDRocket
Sign in

AI agents can't yet do open-ended AI research

AI as Normal Technology Sayash Kapoor

Researchers tested top AI agents on real, unpublished AI research questions—not just tidy benchmarks. Both agent-written papers got flatly rejected by the human experts who posed the questions.

Forget the leaderboards for a second. A team led by Princeton's Sayash Kapoor just ran an experiment that cuts against the industry's favorite narrative about AI agents replacing researchers any day now.

Here's what they did. They found two academics with unpublished AI papers still in progress, and asked them to hand over just the research question—no methods, no data, no answers. Then they gave frontier AI agents thousands of dollars in compute credits and six days to independently answer those questions and write up a paper. The original human authors, who'd already spent months solving the same problem, graded the results blind to nothing—they knew it was AI, but judged the work on its merits.

Both papers got rejected. Not narrowly, not on style points. Unambiguously. And the failure modes, which the team spent over 100 hours dissecting from agent logs, are the interesting part. The agents pitched genuinely promising directions early on—reviewers were impressed—then abandoned them fast, often after a single bad or synthetic dataset spooked them. Neither agent used even half its compute budget despite explicit encouragement to spend it. When their own AI-generated self-reviews flagged real problems, the agents just tacked on caveats instead of rethinking anything. And both essentially locked in their final approach within 24 hours, never meaningfully pivoting again.

This matters because the entire pitch for recursive self-improvement—AI systems bootstrapping AI research and triggering explosive progress—assumes agents can eventually do the messy, ambiguous, ill-defined part of science, not just grind through tasks with a clear pass/fail signal. Kapoor's team, notably skeptics of imminent superintelligence, are careful to flag the study's limits: two papers is a tiny sample, and reviewers knew going in that they were grading a machine. They're calling this method a 'shadow evaluation' and plan to run more.

The bigger point buried in the piece is about bottlenecks. If AI progress runs into several hard, non-verifiable constraints like this one, then Amdahl's law applies: automating the easy 90% barely moves the needle if the remaining 10% is still gated by human-style judgment. Nobody quite knows yet whether open-ended research is one of those permanent speed bumps or just a scaffold problem waiting for a fix.

My take

This is a useful reality check against the breathless RSI hype coming out of the big labs, and it's refreshing that it comes with the researchers' own biases disclosed up front rather than buried. Two papers is a small sample, sure, but the specific failure patterns—no backtracking, ignoring instructions, padding weak findings instead of fixing them—sound a lot less like 'almost there' and a lot more like agents that are still fundamentally following scripts, not doing science. Anyone forecasting an intelligence explosion on the back of benchmark scores alone should be reading the failure logs, not just the leaderboard.

Read more about this at: AI as Normal Technology

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.