TLDRocket
Sign in

AI Handles Incidents, Engineers Lose Touch with Their Systems

Sylvain Kalache

Opinion — commentary, not a factual news event.

AI tools are now fixing incidents on their own. That’s great until humans lose practice and get stuck on the weird ones.

Based on reporting by Sylvain Kalache — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

AI-assisted incident response has crossed from prototype into everyday usefulness. Tools can now inspect alerts, form a theory, pull telemetry, connect recent deployments, and in some cases carry out the fix themselves. The appeal is obvious, especially at 2 a.m. when a routine capacity issue resolves before anyone has to wake up.

But there’s a catch hiding inside the convenience. The more often software handles the boring incidents, the fewer chances engineers get to build the instincts that come from working through them. Then the hard case arrives — the one no automation has seen before — and the humans on call are expected to step in with less practice than they used to have.

The source frames that as the old “Ironies of Automation” problem Lisanne Bainbridge described in 1983: machines take over the routine work, while people remain responsible for the weird, high-stakes failures. Aviation offers the clearest example. Pilots let automation do most of the flying, yet they still train for rare emergencies such as engine failures, unreliable instruments, rejected takeoffs and stalls. Modern turbine engines, the piece notes, see fewer than one in-flight shutdown per 100,000 engine flight hours, which is rare enough that a pilot could go a whole career without one outside a simulator.

Software teams do not get the same discipline by default, so the article argues for building it deliberately. At Rootly, the author says the company worked with Uptime Labs on realistic incident simulations where engineers run a mock e-commerce outage, use observability tools, and coordinate with LLM-powered stakeholders in Slack. The point is not to watch an agent explain itself afterward. It is to practice the messy parts: making sense of incomplete data, keeping the response organized, and dealing with the CEO and customer support while the system is on fire.

That matters even more as LLMs take on more work. The author calls the widening gap between system behavior and human understanding “comprehension debt,” and the fix is old-school on purpose: hands-on control, unfamiliar failures, pressure, and repeated drills. In other words, better automation should come with more training, not less, because the hardest incident is exactly where the humans still have to be good.

My take — AI-written commentary, not fact-checked reporting

This is the bit everyone should stop hand-waving past: automation doesn’t just reduce toil, it erodes touch. The industry loves shipping copilots and then acting surprised when the people are a little rusty. If a team can’t explain the system without asking the system, that’s not maturity, that’s a very expensive trap.

Read more about this at: Sylvain Kalache

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.