Agents have made CI the bottleneck. Faster pipelines are the wrong fix.
The New Stack Arjun Iyer
Agents are flooding CI with more PRs, and the old pipeline is choking. The real bug isn’t speed — it’s that CI still checks a repo, not the whole system.
Based on reporting by The New Stack, Arjun Iyer — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Three September posts point to the same problem from different angles. Anthropic said its CI job volume grew 25x in six months, while its engineers are now shipping about 8x as much code per quarter as they did from 2021 to 2025. Linear said its test suite has nearly quadrupled since January, with agents now writing most of the tests. Depot’s CEO added the broader thesis: CI is changing, and agents need a way to validate code and keep trust while they work.
The immediate reaction has been to make pipelines faster. That means better runners, smarter test selection, bigger caches, and systems agents can call before a commit lands. Those fixes help, and they’re needed. But they all leave one assumption untouched: that the thing being verified is a repository.
That assumption breaks down fast in distributed systems. A repo is one service out of many, not the system itself. A change can pass unit tests, clear CI quickly, and even behave in a branch sandbox, then fail when it hits a real service boundary. The seam is where things go wrong: a renamed field, a tighter timeout, a schema change, an endpoint that works in isolation but not in context. Faster CI doesn’t see any of that if it’s only checking the repo.
Some teams have already started moving verification earlier. Cursor said in February that its agents run in cloud sandboxes, and more than 30% of the PRs it merges now come from agents working that way. GitHub’s Copilot cloud agent runs tests in an ephemeral environment powered by GitHub Actions. Codex uses a setup script and cached containers. Devin boots from environment blueprints. Greptile’s TREX runs the branch and attaches logs and screenshots. But those environments still mostly contain the repo, the branch, and whatever the setup script can install.
The more useful shift is to let the agent verify against the system, not just a copy of its own code. That sounds expensive until you stop cloning entire stacks for every run. A shared Kubernetes cluster can host one stable version of each service and spin up lightweight test environments on top, with only the changed service deployed. The modified service talks to real dependencies, while the rest of the system stays shared and stable. That makes system-level checks cheap enough to run in the agent loop, not after it.
The harder part is governance. Agents need a structured set of approved actions, so they’re not improvising their own tests every time. Platform teams write those steps once, and the output becomes an artifact review tools can read: what was sent, what was touched, what held. CI still matters, but it should confirm a change already exercised against live services, not be the first place cross-service breakage shows up.
My take — AI-written commentary, not fact-checked reporting
The industry keeps trying to speed up the wrong checkpoint. Faster CI is nice, but it’s still a very expensive way to discover that a distributed system is not a repo with delusions of grandeur. The real competition is who can give agents safe access to the actual system without turning every change into a small fire drill.
Read more about this at: The New Stack