TLDRocket
Sign in

[AINews] Lilian Weng summarizes 35 papers on Harness Engineering for RSI

Latent Space

Lilian Weng dropped a big research recap on how AI agent "harnesses" - the scaffolding around models - drive self-improvement, not just raw model weights. It's a signal that even top researchers see agent design, not just bigger models, as the next frontier.

Based on reporting by Latent Space — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Lilian Weng, now a cofounder at Thinky, spent her latest post walking through the research literature on what she calls harness engineering - the scaffolding, tooling, and workflow design that sits around a model and shapes how it acts. Her framing is notable: recursive self-improvement, she argues, doesn't have to mean a model rewriting its own weights. It can mean the harness around it getting smarter. And even when harness tricks eventually get folded into the core model itself, she notes, the underlying job of specifying goals and context for that model never actually goes away.

The post surveys proven design patterns and recaps optimization work ranging from the well-known ACE paper to newer Meta-Harness approaches that Latent Space's AINews has been tracking anecdotally for a while. Weng isn't alone in pushing this idea. Greg Brockman has been quietly signaling something similar, and now a respected researcher running her own lab is saying it out loud, which gives the harness-first view of agent progress a lot more weight than a single hot take would.

Elsewhere in the ecosystem, the harness obsession showed up everywhere at once. Anthropic pushed its Claude Cowork background-agent product to mobile and web, treating Claude less like a chat window and more like a teammate quietly running tasks in the background. LangChain launched a Deep Agents course and an open-source harness project. Google rolled managed background execution, remote MCP servers, and credential refresh into its Gemini API. Even smaller infrastructure moves, like Weaviate's runtime-gated write access for its MCP server or Hermes Agent's pluggable secrets managers, point the same direction: the interesting engineering right now is happening in the scaffolding, not necessarily inside the model weights themselves.

Meta, meanwhile, made noise on the model side. Muse Image and the previewed Muse Video landed with an explicitly agentic generation loop - planning, web search, tool use, code execution, self-refinement - before an image or clip ever gets rendered. Meta says that self-refinement behavior wasn't hand-scripted; it emerged during reinforcement learning. Muse Image quickly climbed to number two on Image Arena, trailing GPT Image 2, while Muse Video debuted at number three on Video Arena. No technical paper has surfaced yet, which is a strange gap for a launch this prominent, but the agentic loop itself fits neatly into the same story Weng is telling: increasingly, what a model does with its harness matters as much as what the model knows on its own.

My take — AI-written commentary, not fact-checked reporting

Weng putting her name on the harness-first view of self-improvement should settle an argument that's been simmering for months: scaffolding isn't a workaround for weak models, it's where a lot of the real engineering now lives. Companies still chasing bigger weights while ignoring how agents plan, retry, and call tools are optimizing the wrong layer. Meta shipping an agentic image loop with zero technical documentation is exactly the kind of move that will age badly once someone actually checks whether the self-refinement claims hold up outside a leaderboard.

Read more about this at: Latent Space

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.