Reading today's open-closed performance gap
Interconnects Nathan Lambert
Open AI models are catching closed ones on coding tests. But the real edge may be shifting to messier, unmeasured tasks.
Based on reporting by Interconnects, Nathan Lambert — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Every few months someone publishes a chart claiming to measure exactly how far open-weight AI models trail the closed frontier, and every few months that chart tells you less than it seems to. The most cited version is Artificial Analysis's Intelligence Index, a roll-up of roughly ten sub-benchmarks meant to track the state of the art. The number moves, headlines get written, and the actual texture of what's being measured — and how fast that thing changes underneath everyone's feet — gets lost.
Look at the arc since ChatGPT launched. The early focus was chat quality, basic math, and simple coding, propped up by instruction tuning and RLHF. That saturated fast. By 2025, with reasoning models as the default, the action moved to harder coding tasks and light agentic work, all powered by reinforcement learning with verifiable rewards. Gemini 3 posts staggering scores on these benchmarks right now, yet barely shows up in the agentic tools people are actually paying for and deploying daily. That mismatch alone should worry anyone treating a leaderboard as ground truth.
The next fight is over messier ground: accounting, law, healthcare, and other knowledge-work domains that require real integrations, not just clean verifiable answers. Frontier labs in the US are pouring money into building the training environments and datasets for this next wave, and — this is the part people miss — fast-following labs, many in China, aren't just distilling outputs from bigger models. They're buying or replicating the same specialized environments at a discount once they exist. Chasing benchmarks is genuinely achievable for them precisely because benchmarks, by definition, are things you can build an environment to mirror.
That creates real economic pressure on OpenAI and Anthropic. Their enterprise revenue leans heavily on being clearly better at coding and terminal-based agent work. If that gap narrows — and cheaper open alternatives become swappable — the premium pricing has to be justified by something other than raw capability: sticky customer relationships, better tooling, habit. So the incumbents need to keep finding new frontiers to own, over and over, just to keep the growth story intact.
None of this means open models have caught up everywhere. On oddball evaluations like WeirdML or ARC-AGI-2, the gap is still wide, and daily use exposes real weaknesses — shakier long-context handling, more frequent need to reset an agent's memory than you'd get with Claude or Codex. But calling the leading Chinese models mere benchmark-chasers, or reducing their gains purely to distillation, ignores how much of their progress comes from smart use of RL environments. They're closer to the frontier than most predictions from a year ago allowed for, and the interesting question now isn't whether they're catching up — it's on which specific ground they will, and won't.
My take — AI-written commentary, not fact-checked reporting
I've long thought the open-vs-closed debate gets flattened into a single scoreboard because that's easier to tweet than admit the frontier keeps mutating underneath us. What this piece nails is that closed labs aren't winning on model quality alone anymore, they're winning on knowing which goalposts to move next, and that's a much shakier moat than the hype cycle wants to admit. If your business case for a $100B compute buildout depends on perpetually inventing new benchmarks nobody else can afford to chase, you don't have a technology advantage, you have a spending advantage, and those expire the moment someone else's GPU bill catches up.
Read more about this at: Interconnects
Related stories
Import AI 465: Open vs closed gaps; Kimi K3; Demis' big policy plan
Import AI · 1 month ago ·
20