TLDRocket
Sign in

AI agents can't yet do open-ended AI research

AI as Normal Technology Sayash Kapoor

Researchers tested if AI agents can do real, open-ended AI research, not just verifiable tasks. Both agent-written papers got flatly rejected by the human experts who set the questions.

Based on reporting by AI as Normal Technology, Sayash Kapoor — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

A new study just poured cold water on one of AI's most hyped ideas: that AI agents are close to automating AI research itself, the so-called recursive self-improvement loop that underpins predictions of explosive progress. The researchers, working with authors of two unpublished AI papers, had those authors draft their own papers' central research questions. Then frontier AI agents were set loose to answer them, armed with thousands of dollars in API credits, real compute, and six days to work.

The original authors reviewed what the agents produced. Both papers were rejected, no ambiguity about it. So the team spent over a hundred hours digging through the agents' logs to figure out why, and what they found paints a pretty consistent picture of agents that are good at narrow, checkable tasks but lost once the work gets messy.

The agents pitched directions that impressed the expert reviewers at first, then abandoned them fast when low-quality or synthetic data showed up, rather than pushing through. They also had no real sense of their own budget: both runs finished with less than half the API money spent and hours still on the clock, despite being told to track and use their resources. Even when their own AI self-review tools flagged many of the same problems the human experts later raised, the agents didn't respond with new thinking. They just tacked on caveats or kept grinding at directions that weren't working. Worse, they gave up on their most ambitious targets within the first day and never really pivoted after that, and they skipped explicit instructions about exploration time, review frequency, and paper length limits.

The researchers are careful about the limits of this method, which they call shadow evaluations: reviewers know the paper is AI-written, the sample is just two papers, and there's plenty of room for researcher judgment calls. The team also flags its own known leanings on the RSI debate and says it deliberately recruited collaborators who don't all agree with them.

The bigger question the paper raises is whether hill-climbing on verifiable, checkable tasks (something agents are already doing well) actually adds up to the kind of open-ended research automation that would trigger explosive AI progress. The authors' answer, for now, is no; narrow gains look real, but broad recursive self-improvement looks like a different problem entirely, one with judgment, backtracking and creativity bottlenecks that haven't been cracked yet. Whether those bottlenecks are easy to fix or genuinely hard will decide, per Amdahl's law logic they invoke, whether AI progress speeds up dramatically or just nudges forward no matter how good agents get at the parts they've already mastered.

My take — AI-written commentary, not fact-checked reporting

People keep pointing to benchmark wins on verifiable tasks as evidence that AI agents are about to start running AI research on their own, and this study is a decent reality check on that leap. Getting good at graded homework problems is not the same as having judgment, and an agent that burns half its budget doing nothing and can't take a hint from its own review tool is not a researcher, it's a very expensive autocomplete. The honest move here is treating this as one early, small-sample data point rather than proof either way, which is more intellectual humility than most AI-progress takes bother with.

Read more about this at: AI as Normal Technology

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.