TLDRocket
Sign in

Agents keep changing their answers. Harness just built delivery pipelines that don’t care.

The New Stack Adrian Bridgwater

Harness built pipelines to test and govern AI agents like regular code, with quality gates instead of trust. Only 17% of firms use agents so far—turns out unpredictable answers make deployment scary.

Based on reporting by The New Stack, Adrian Bridgwater — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Barely one in five organizations has actually put an AI agent into production, according to Gartner's 2026 CIO survey. Harness, the software delivery lifecycle company, thinks it knows why: agents don't behave like code, and nobody's built the guardrails to treat them that way. This week the company launched its AI Agent Development Lifecycle service, aiming to run agents through the same governance, testing, and security checks that application code already gets before it ships.

Trevor Stuart, SVP and general manager at Harness, frames the problem bluntly. Getting an agent to work in a demo is easy. Knowing it'll behave the same way in production is a different animal entirely, because agents aren't deterministic the way code is. Feed the same input to the same agent twice and it might pick a different tool, take a different path, produce a different answer. A test that passes today tells you nothing about tomorrow. That breaks the standard playbook for catching bugs, since incidents stop being reproducible on demand. The result, Stuart says, is that companies keep agents locked in sandboxes indefinitely — technically a success, practically useless.

Harness's answer isn't to make agents predictable. It's to make the pipeline around them predictable. Instead of checking whether tests passed, Stuart wants teams grading responses on correctness, safety, and performance, then wiring that score into the pipeline as a straightforward pass-fail gate. Nobody's claiming agent output becomes reproducible anytime soon. But the record of what an agent did — every model call, every tool call, every step — can be captured and made reproducible, which at least gives engineers something concrete to tune against instead of guessing.

Five new products carry that philosophy. AI Evals lets teams define datasets and scoring functions so quality regressions get caught automatically. Agent deployments extends the canary releases and approval guardrails Harness already runs for Kubernetes onto managed agent runtimes. AI configs handles prompt and model changes at runtime. The AI asset catalog scans an org's repositories to find every agent, skill, and plugin and tie it to an owner. And AgentTrace records what happens during a run — where an agent slowed down, which path it took, how swapping a model or prompt changes the outcome. Harness is open-sourcing the tracing components behind that last one, harness-sdk and harness-evals, so other developers can bolt the same instrumentation onto their own tools.

Stuart won't pretend there's a tidy ROI figure attached to any of this. Harness's own 2026 State of Engineering Excellence report found that 31% of a developer's day goes to AI-related work that shows up in no metric anywhere, and 94% of engineering leaders admit their tracking misses things like tech debt and burnout entirely. What the asset catalog is meant to fix, at least, is accountability: ownership attaches the moment something gets built rather than sitting with some vague governance role nobody can find.

The timing lines up with Harness's June 2026 launch of Autonomous Worker Agents, which let teams run agents as governed steps inside existing delivery pipelines. Agent DLC just stretches that same governance across the agent's entire lifecycle — eval gates, deployment approvals, security checks, all running as stages in one pipeline from creation onward.

My take — AI-written commentary, not fact-checked reporting

Seventeen percent adoption isn't a rounding error, it's an industry quietly admitting agents aren't ready for the room they're being pushed into. Harness deserves credit for refusing the easy sell — nobody here is promising agents will suddenly behave, just that the mess around them can finally be tracked and owned. That's a more honest pitch than most of the agent hype circulating right now, and probably the only one that survives contact with an actual production incident.”}

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.