TLDRocket
Sign in

The Evolution of the Agent Harness

Latent Space Dan McAteer Covered by 4 sources

AI agents finally started working around Christmas 2025, and the reason may be the model and the harness hitting their stride together. That shift matters because the real bottleneck is moving from model smarts to human attention.

Based on reporting by Latent Space, Dan McAteer — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Sometime around Christmas 2025, AI engineers noticed something awkwardly simple: agents started to work. Not perfectly, not everywhere, but well enough to feel like a step change. The hard part is pinning down why. Maybe people finally had time over the holidays to try the newest agents on the newest models. Maybe the models had crossed a threshold. Maybe the wrappers around them had gotten better. The cleaner answer is that both curves moved, and they crossed at the right moment.

That is the core argument here: the jump wasn’t just in the weights. It was in the system built around them. The article points back to Lukasz Kaiser’s June comments on “Unsupervised Learning,” where he described last winter’s change as hard to isolate because the harness changed, post-training changed, and new pre-trained models arrived together. In that reading, the harness is everything besides the model itself: the tools, environment, context, memory, permissions, and guardrails that let the model act in the real digital world instead of sitting in a prompt like a brain in a vat.

The history comes in stages. ReAct, from October 2022, was the “agent loop” on paper: reason, act, observe, repeat. Toolformer, in February 2023, hinted that tool use could be trained rather than just prompted. Then came AutoGPT and BabyAGI, where the harness outran the models and gave them too much autonomy. The result was predictable enough to be depressing: a loop cannot create capability that isn’t there. It only magnifies what exists, and below a certain threshold it magnifies failure.

By 2023 and 2024, tools like Cursor and Copilot pulled back from that mistake by keeping the human in the loop. And near the end of 2024, o1 changed the equation again. For the first time, the model side looked strong enough that the harness was the thing holding it back. Claude Code, launched in February 2025, seized that moment by giving the model bash plus file read and write access, then relying on permission rules instead of asking a human to approve every move. The article says that timing mattered as much as the product itself, and that Claude Code reached roughly $1B ARR within six months because Anthropic hit the crossover point.

Now the braid is tighter. Harness-Bench showed the same model scoring from 52.4 to 76.2 across 106 tasks just by changing the harness. OpenAI reported a similar effect on ARC-AGI-3, where GPT-5.6 Sol’s score rose from 13.3% to 38.3% after adding retained reasoning and compaction. The pattern is clear enough: reinforcement learning is moving into the harness, and then the model starts absorbing the harness back into its weights. OpenAI’s codex-1 was trained with reinforcement learning on real-world coding tasks in different environments. GPT-5.1-Codex-Max was described as the first model natively trained to operate across multiple context windows through compaction. Anthropic’s Thariq Shihipar said the team deleted 80% of Claude Code’s system prompt. Production by reduction, basically.

The punchline is that the next bottleneck is probably not more model intelligence. It is human attention. The article’s bet is that the next essential layer will be a human attention policy surface: when the agent may interrupt, when it should keep going, what it can decide alone, and what needs approval. The harness used to be the model’s body. If the model keeps absorbing the body, what’s left is the interface to the one thing still in short supply.

My take — AI-written commentary, not fact-checked reporting

This is the part people keep ducking: if the model can absorb the whole stack, the scarce resource doesn’t vanish, it moves upstairs to the human. That’s why the next real product category is not another agent brag sheet, but a way to say “stop pinging me unless it matters.” The AI industry has spent years trying to automate away friction; now it gets to automate the nuisance instead.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.