TLDRocket
Sign in

The Sequence Knowledge - Issue 937: RSI in Post-Training: The Loop That Already Shipped

TheSequence Jesus Rodriguez

The real self-improvement loop in AI is post-training, not code-writing. Labs have been running it at scale for two years, and it’s already what made frontier models into agents.

Based on reporting by TheSequence, Jesus Rodriguez — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Forget the fantasy of an AI rewriting its own source code. The loop that matters most is much duller, and much more useful: a model makes several candidate answers, something scores them, the good ones get turned into training data, and the model learns to do that part better next time. Then it happens again. And again.

That loop has been sitting in the middle of frontier AI for two years at industrial scale, even if it gets less attention than the shiny talk about self-improving agents. The Sequence’s point is simple: the factory got automated before the design office. In AI, the line between those two depends on whether the task comes with an answer key.

The naming has changed, not the mechanism. A 2022 paper called it STaR. In 2025, the same idea is showing up as RLVR. Inside companies, it’s just post-training. Same machine, different label.

And this is not a side detail. The source argues that this is the loop that took frontier models from chatbots to agents. That makes the whole RSI debate look a bit backwards: the self-improvement that already ships is not the dramatic code-rewriting version. It’s the quieter one, built into the training pipeline, churning through answers and feeding on its own best outputs.

My take — AI-written commentary, not fact-checked reporting

People keep waiting for the dramatic robot breakthrough, while the boring pipeline keeps doing the real work. That’s usually how tech changes the world: not with a grand speech, but with a grading loop and a spreadsheet. The obsession with flashy self-modification is a nice distraction for the demos.

Read more about this at: TheSequence

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.