Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
MarkTechPost Asif Razzaq
Google and top university labs open-sourced RRSI, a way for AI agents to improve their own harness without changing model weights. It aims to stop the usual overfitting trap, so gains stick on tasks the agent never tuned for.
Based on reporting by MarkTechPost, Asif Razzaq — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Google Cloud AI Research has released RRSI, short for Regularized Recursive Self-Improvement, with researchers from UNC-Chapel Hill, Stanford, and Washington University in St. Louis. The idea is unusual but tidy: let an LLM agent rewrite the scaffolding around itself — prompts, tools, memory, control flow, even sub-agents — while the model weights stay frozen.
That choice matters because a lot of self-improvement loops can get lazy. They keep proposing edits, scoring them on the same evolve set, and then picking the winner. Do that often enough and the system can start learning the test rather than the job. The RRSI paper calls out three ways this goes wrong: benchmark-specific fitting, noise chasing, and complexity piling up until the evolve score and real transfer drift apart.
RRSI tries to regularize the search itself. On the proposal side, it uses an annealed edit budget so early rounds can bundle several edits, while later rounds are limited to a single attributable change. It keeps a ledger of each candidate’s component, hypothesis, diff, score change, and cost change, and the proposer reads that history so false leads are less likely to come back. When progress stalls inside the noise band, exploration shifts to components the run hasn’t touched yet.
On the selection side, there’s a leakage critic that blocks task names, entities, answers, or benchmark-specific logic before scoring. Gains also have to clear a noise-adjusted floor measured on the unchanged base harness, and any extra inference tokens have to earn their keep. If a component stops producing gains, pruning can make it a deletion target. The research team explicitly compares those rules to classic regularizers: the edit budget to L0, pruning to Lasso, and the cost rule to Ridge.
The paper reports improvements across eight benchmarks. Terminal-Bench 2.1 rose from 74.2% to 80.2% on the evolve split, SWE-bench Verified climbed from 82.0% to 83.8% on a split never used for selection, and held-out sets like JobBench, GDPval, and APEX-Agents also moved up. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 went from 64.6 to 78.7, while SWE-bench Verified moved from 76.8 to 79.0.
There’s also a cost angle. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial, versus 3.80M for unregularized evolution. And in a comparison table, it’s the only method with an out-of-distribution average more than 1 point above H0, which is the sort of boring, useful result that self-improvement papers usually skip past.
My take — AI-written commentary, not fact-checked reporting
This is the right kind of self-improvement story: make the harness smarter, don’t let the model doodle on its own brain. The industry loves loops that impress the evolve set and then embarrass themselves in the wild; RRSI is basically a small rebellion against that habit. More of that, less magical thinking dressed up as autonomy.
Read more about this at: MarkTechPost