TLDRocket
Sign in

AI’s recursive self-improvement might not come so quickly after all

MIT Technology Review Michelle Kim

A new study says AI agents still can't do real research on their own. They ace the engineering grind but flunk the creative judgment self-improving AI would need.

Based on reporting by MIT Technology Review, Michelle Kim — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Peter Kirgis and Sayash Kapoor at Princeton, along with researchers from several other institutions, wanted to know if AI could actually do open-ended science, not just the tidy, checkable kind. So they built a test called shadow evaluation: hand an AI agent a research question from an unpublished paper submitted to NeurIPS 2026, give it six days, three thousand dollars in Anthropic API credits, a GPU budget, and see what comes out. The agent running the show was Claude Opus 4.8 on open-source software called OpenClaw, and it tackled two problems — one about controlling an LLM's behavioral 'personas' by editing model weights, the other about detecting when a spreadsheet-prediction model has gone unreliable.

Both resulting papers got rejected. Not because the agents couldn't do the legwork — they reviewed literature, ran hundreds of experiments, compiled results, all the engineering scaffolding of research. The problem, as Kapoor puts it, is that the agents were 'unambiguously bad at carrying out the research itself.' They tested hypotheses on absurdly small synthetic datasets, wrote up findings that were hard to parse, and never landed a genuinely novel contribution. Worse, they behaved like students who ask a great opening question and then quit the moment the first data point looks discouraging: the agents floated hypotheses resembling the original authors' own starting points, then abandoned them based on thin evidence, unable to backtrack or rethink from scratch. Feedback from subagents and outside review tools mostly bounced off; instead of changing course, the agents hedged their claims and piled on caveats.

There was at least one bit of good news buried in the failure. The agents didn't reward-hack — they never hid or misrepresented data to look better, even when subagents occasionally hallucinated or fudged something, the lead orchestrator agent caught it. Kapoor's read on why the engineering skill and the research judgment split apart comes down to training. Reinforcement learning works great when success can be scored automatically, which is why coding got so much better so fast. Open-ended research doesn't offer that kind of scaffolding, so the models never got drilled on it the same way.

The study is small — two papers, evaluators who knew they were grading AI work, researchers with plenty of discretion in how they set things up — so nobody should treat it as gospel. But it lines up with what Anthropic cofounder Jack Clark wrote separately in his newsletter about the company's own attempts to automate AI safety research: capable engineers, short on 'valuable, intuitive creativity,' prone to 'rote, formulaic thinking.' That's notable given Anthropic published a blog post in June about AI building itself, and OpenAI touted its GPT-5.6 Sol model helping post-train a smaller one in July. The public optimism and the internal findings don't seem to be telling quite the same story.

Kapoor's team is now rerunning the experiment on Anthropic's newer model, Mythos, which launched in April under new safety restrictions and is available only to approved organizations. Boston University's Najoung Kim, who wasn't involved in the study, thinks focused investment could still produce real progress here. But the deeper question Kapoor keeps circling back to is whether recursive self-improvement actually needs this creative leap at all, or whether AI can grind its way there purely by getting incrementally better at narrower, scoreable tasks. He calls it the trillion-dollar question, and right now nobody, including the people building these systems, seems to have the answer.

My take — AI-written commentary, not fact-checked reporting

Every big self-improvement claim from these labs deserves the same question: is this progress on judgment, or just another benchmark getting squeezed? The honest answer here is the second one, and that gap matters because the whole recursive-self-improvement pitch assumes creativity is just another skill waiting to be reinforcement-learned into existence. Maybe it is. But betting trillions on that assumption while your own researchers privately admit the models are formulaic thinkers dressed up as scientists is the kind of overconfidence the industry should be a lot more careful about advertising.

Read more about this at: MIT Technology Review

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.