AI agents are creating more work, not less — and OpenAI’s own numbers back it up
The New Stack Amanda Caswell ● Covered by 5 sources
OpenAI says its research agents are now doing more work than people across the org. That sounds like progress, until you see how much human supervision and spend still pile up.
Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
OpenAI says it has hit the target it set last fall: researchers are now using what it calls an “automated research intern,” an agent built to take on well-defined tasks that would otherwise keep a researcher busy for days. By mid-August, the company says its agents were logging 3.1 agent-workdays for every human workday across the research organization. That’s a lot of machine activity. It is not, by itself, proof of a lot of finished science.
The company’s numbers also show how expensive this gets. The median researcher was spending more than $600 a day on inference at API prices, while the 90th percentile was above $7,000. And as coding-agent use climbed through 2026, OpenAI says the work spread across six buckets from Decide to Communicate, with activity rising in all of them between January and August. The least automated part was still the most human one: deciding what research to pursue.
This is the catch with agent metrics. OpenAI converts agent time into standard eight-hour workdays, but several agents can run at once, so the figure says more about how long the software was busy than what it actually finished. The company says code output and experiment counts are easy enough to measure, but neither tells the full story of progress. In fact, even on tasks that took four to eight hours, humans still had to step in on more than half of the ones agents managed to complete successfully between January and July.
The practical strain is already showing up. OpenAI says agents are writing research and infrastructure code, monitoring experiments, and helping enough that one team stopped holding debugging office hours altogether. But more agent hours also mean more code to review, more runs to check, and more chances for something to drift off course. On July 20, agent-caused outages were bad enough that OpenAI took its training container service offline, then brought it back with tighter restrictions. On August 7, it tightened access again after early signs Astra could hit the “Critical” level in its Preparedness Framework, limiting it to higher-security research areas.
And when Astra usage fell 59.2% the next week, the work did not disappear. Other models picked it up, with GPU allocation rising 17.2% and covering about 85% of the drop. OpenAI has a clear next milestone in mind — an automated AI researcher by March 2028 — but for now the real story is simpler: the agents are multiplying the work, and the humans are still stuck doing the cleanup.
My take — AI-written commentary, not fact-checked reporting
This is classic AI theater: the demo looks like automation, the back office looks like supervision with a higher cloud bill. OpenAI keeps proving the same uncomfortable point that applies across the whole agent craze — you don’t delete work, you move it around and call it progress. Sometimes the most advanced thing a model can do is create a fresh queue for a human to babysit.
Read more about this at: The New Stack