TLDRocket
Sign in

Today’s Codex will feel “primitive” by fall — and its own team’s roadmap backs it up

The New Stack Amanda Caswell

OpenAI's own product lead says Codex will look primitive in 2-3 months. Reason: future models need cloud muscle, not your laptop.

Thibault Sottiaux doesn't run some rando dev shop — he leads core products at OpenAI, so when he posted on X late Monday that Codex is "a good harness" but will feel "primitive in 2-3 months," people paid attention. His follow-up line was the real tell: "the next generation of models need more than your laptop." No roadmap, no dates, just a warning shot from inside the building.

The timing lines up with something OpenAI already announced. Back in June, the company said it plans to buy Ona, formerly known as Gitpod, a firm that builds secure cloud dev environments already used by 2 million developers. OpenAI framed the deal as the "next phase of Codex" — one where an agent keeps grinding away in a customer's cloud long after the laptop that kicked off the task gets closed and stuffed in a bag. Right now, Codex leans on cloud compute but still often needs that laptop nearby to reach a project's files and tools. Kill the connection, and the agent can lose its footing mid-task.

The ambition here isn't small. OpenAI showed in a February test that Codex could run solo for about 25 hours, burn through 13 million tokens, and spit out roughly 30,000 lines of code building a design tool from nothing. Impressive, until you clock what Alibaba just did: its Qwen3.8-Max agent worked autonomously for 16 straight days, landing 265 commits without a single human nudge. That's the bar apparently being set for late 2026, and it's the kind of endurance no laptop session survives.

Getting there means solving a much messier problem than model quality. An agent with standing access to a company's network, credentials, and CI/CD pipeline is a genuinely scary attack surface, and OpenAI knows it — hence the promise that Ona's setup stays customer-controlled, with OpenAI supplying only the model and orchestration layer. Anthropic is chasing a similar itch with its Mendral acqui-hire, aimed at automating flaky-test cleanup and dependency reviews. Everyone building these persistent agent environments is quietly rediscovering that autonomous coders need identities, audit logs, and access rules — basically the same governance headaches human engineers already have, just handed to software that never clocks out.

So Sottiaux's prediction isn't really hype about a smarter model dropping soon. It's an admission that Codex's current shape — bound to a session, tethered to a machine — is already the bottleneck, and OpenAI is racing to build the infrastructure that removes it before someone else does.

My take

Calling your own two-month-old product "primitive" isn't candor, it's marketing dressed up as humility — a neat trick for keeping developers hooked on the upgrade treadmill instead of settling into whatever they're using today. The actual story worth watching isn't model horsepower, it's who builds the boring plumbing — identity, access control, audit trails — for agents that run unsupervised for days. Whoever nails that governance layer quietly, without a viral tweet, ends up owning the market.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.