The twilight of the chatbots
One Useful Thing Ethan Mollick ● Covered by 3 sources
AI models are now finishing multi-week engineering jobs in hours, and how people work with them is quietly flipping from chatting to managing agents. The leap from prompting a chatbot to overseeing swarms of agents is happening faster than most workplace plans can keep up.
Based on reporting by One Useful Thing, Ethan Mollick — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Ethan Mollick's latest dispatch is less about a single model release and more about the shape of the curve itself. Benchmarks like METR, GDPval, and the UK AI Security Institute's tests are all showing the same thing: the amount of useful human-work-hours an AI can complete from one prompt is climbing at a better-than-exponential clip. Epoch's headline number is the kind of thing that should stop you mid-scroll — Anthropic's Opus 4.7 ran unsupervised for 14 hours and produced a software package that would normally take a human team two to seventeen weeks, for $251 in token costs. Mollick says his own tests with Fable produced similarly startling results: nine hours of autonomous work replacing what a team would need over a week to finish.
There's a second track running behind the American frontier labs — Anthropic, OpenAI, and a quieter Google — and it's made up entirely of Chinese open-weights models. They trail six to twelve months behind the proprietary leaders but are climbing their own exponential curve, and because they're open and cheap to run, that gap matters less than the raw numbers suggest. Mollick's harbor-simulation test, where different models had to build an evolving interactive scene, is a nice reminder that benchmarks flatten out real differences in taste, judgment, and design instinct that only show up when you actually watch the things work.
The more consequential shift, though, is behavioral rather than technical. The old model of AI use — prompt, check, prompt again, basically co-piloting — is losing ground to something closer to management. Long-running agents with tool access and dedicated harnesses, things like Claude Code or Codex, don't need a human hovering over every step. OpenAI's internal data, gathered with academic economists, shows this happening company-wide, and not just among engineers: legal and HR teams are adopting agents at nearly the same pace as coders. A quarter of OpenAI staff are now juggling four or more agents running simultaneously each week.
What's interesting is who actually gets good results out of these systems. A study of Claude Code users found that coding skill mattered less than domain expertise — the more someone knew about their own field, the better they were at directing an agent toward useful output, regardless of whether they could code at all. That flips the old assumption that chatbots mainly help non-experts paper over their gaps. Now it looks like experts, treating agents as employees rather than search engines, are the ones pulling ahead.
Mollick's broader point is about perception, not just capability. Any AI plan drafted before late 2025 assumed a system good for a couple of error-prone hours of work; a few months later, that same system is doing sixteen-plus hours reliably. Because doubling curves feel like sudden jumps to humans who experience time linearly, every policy lurch — a government yanking access to a model, markets repricing entire industries overnight — looks like chaos when it's really just what an exponential looks like from inside the room.
My take — AI-written commentary, not fact-checked reporting
I'll admit the Opus-4.7-builds-two-months-of-software-for-$251 stat is the kind of thing that should worry anyone still pricing engineering headcount like it's 2023. But the part everyone's underrating is the shift from chatbot to manager — that's the actual labor story, not another benchmark chart. Open-weights Chinese models chasing the frontier at a discount is the geopolitical subplot nobody in Washington seems to be pricing in either, and pretending export controls freeze that gap for long is wishful thinking dressed up as policy.
Read more about this at: One Useful Thing