The future of business, today.
RSSNewslettersAdvertise
Business Future Today

Operational resilience

AI Can Resolve More Incidents—But It May Leave Engineers Less Ready for the Hard Ones

AI-assisted operations can shorten routine outages. The management challenge is preventing automation from eroding the operational judgment needed when novel failures break through.

Editorial image for AI Can Resolve More Incidents—But It May Leave Engineers Less Ready for the Hard Ones
Illustration: Business Future Today

AI is moving deeper into incident response: inspecting alerts, querying telemetry, correlating changes and, in some cases, carrying out fixes. For technology leaders, the immediate appeal is straightforward: fewer overnight pages, faster recovery from familiar problems and a lower operational burden on already-stretched teams.

But there is a less comfortable consequence. If automation handles the routine cases through which engineers learn how a system behaves under stress, teams may be less capable when an unusual, high-severity failure requires human judgment.

The automation paradox reaches SRE

In an essay published on his site, Rootly executive Sylvain Kalache argues that AI incident-response systems can create a form of “comprehension debt”: a widening gap between the complexity of production systems and responders’ practical understanding of them.

The concern follows a well-established human-factors idea. In *The Ironies of Automation*, researcher Lisanne Bainbridge described how automation removes much of the routine work that helps operators build experience, while leaving people responsible for the abnormal situations automation cannot handle.

Applied to software operations, the likely result is not uniformly worse performance. Routine mean time to recovery may fall as AI handles known capacity, configuration or deployment-related issues. The operational risk lies in the tail: novel, ambiguous incidents where signals conflict, the blast radius is unclear and the correct response depends on context the system has not seen before.

Those are precisely the moments when a team needs people who can form and discard hypotheses, understand dependencies, make decisions with incomplete information and keep stakeholders aligned.

Faster remediation is not the same as resilience

Executives should avoid treating autonomous remediation as a simple headcount or on-call-cost reduction exercise. It changes the work rather than eliminating accountability.

An AI agent may provide a strong investigation trail and a plausible fix, but responders still need to know when to trust it, when to stop it, and how to recover if its action makes an incident worse. That calls for operational controls: scoped permissions, approval gates for higher-risk changes, clear audit logs and tested rollback paths.

It also calls for a different set of performance measures. Teams that focus only on aggregate MTTR can miss a worsening ability to handle rare but consequential events. Track the share of incidents resolved autonomously, but pair it with measures such as time to establish an accurate incident hypothesis in simulations, quality of incident communications, and the frequency with which responders lead complex exercises themselves.

Treat simulations as a core operating practice

The aviation comparison is useful not because software incidents are identical to flight emergencies, but because pilots rehearse rare failures that they may never encounter in regular operations. Production engineers need a comparable practice loop.

That means running realistic game days, chaos experiments and incident simulations that include both technical and organizational pressure. A useful exercise should force participants to inspect real or representative telemetry, evaluate misleading clues, coordinate roles and communicate with customer-facing and executive stakeholders—not simply follow a prewritten runbook.

AI can support this training. Teams can use an agent to generate plausible failure scenarios, play simulated stakeholders, summarize investigation steps or explain the evidence behind a recommendation. But passive observation is not a substitute for having an engineer run the response.

What to watch next

As AI operations products mature, buyers should look beyond demonstrations of autonomous ticket closure. The stronger question is whether a platform helps preserve human system knowledge: Can engineers inspect the agent’s reasoning and evidence? Can they safely take control? Can the organization turn real incidents into repeatable training scenarios?

The goal is not to keep humans busy with toil that software can remove. It is to use the time AI saves to build the judgment required for the failures no model, runbook or dashboard can fully anticipate.

Sources

STAY AHEAD

The future of business, in your inbox.

Useful signals on the companies, technologies and shifts changing business.

One useful briefing. Unsubscribe any time.