Join live: What happens when your AI goes down?
Join live: What happens when your AI goes down?
TL;DR: Your AIOps tool groups alerts into correlated incidents every week. Your mean time to resolution (MTTR) hasn't moved in months. Both things are true, and that's the problem. When correlation works but MTTR stays flat, post-mortems still run 60 to 90 minutes, and your rotation carries all the setup work, you've hit the AIOps ceiling. Correlation without coordination can't close the incident lifecycle gap. Seven diagnostic signals tell you whether you need configuration tuning or a different tool class. Ours is the second kind: our Investigations triages alerts, analyzes root cause, then drafts the fix as a pull request your team reviews and merges.
Your AIOps (Artificial Intelligence for IT Operations) deployment isn't broken. It filters alert noise, correlates related events, and surfaces patterns faster than manual triage. So why does your MTTR plateau, and why are post-mortems still an archaeology project?
AIOps reduces noise, not resolution time. When alert correlation succeeds but incidents still take just as long to resolve, you're not dealing with a configuration problem. You're dealing with a category fit problem.
AIOps platforms analyze telemetry and event streams to turn raw data into patterns a human can act on. In practice, that means AIOps excels at noise reduction, pattern recognition, and anomaly detection.
The catch: AIOps identifies and groups the signals. It doesn't assemble your incident team, capture your timeline, or coordinate your response workflow.
Alert filtering works. AIOps platforms use clustering and deduplication to reduce alerts to coherent incidents, turning large volumes of raw events into actionable incidents with context attached. Across its 2025 AIOps deployments, Thoughtworks measured L1 and L2 ticket volume falling 35 to 40%, with root-cause analysis cycles shortened from hours to minutes.
But noise reduction and resolution speed are different problems. You can group hundreds of alerts into a dozen incidents and still spend 10 to 15 minutes assembling your team before anyone touches the problem.
As Gartner defines it, AIOps combines big data and machine learning to automate IT operations processes, including event correlation, anomaly detection, and causality determination. The technology works when the problem is pattern recognition at scale.
The limitation emerges when the problem shifts from "what's happening" to "who's fixing it and how do we document what we tried." 68% of site reliability engineers say the scripts still need review because model training lacks context for proprietary middleware, which means you spend the same time interpreting and validating what the AI suggests.
AIOps and deterministic scripts solve different problems. Check which kind yours is.
| Approach | Best for | Limitation | Example |
|---|---|---|---|
| AIOps pattern recognition | Unknown failure modes, complex system interactions, high-cardinality data | Requires continuous tuning | Correlating hundreds of alerts from dozens of microservices into root incidents |
| Deterministic scripts | Known failure modes, repeatable processes, straightforward remediation | Limited to predefined scenarios | Automated restart of a service when memory exceeds defined thresholds |
The AIOps ceiling is the point where tuning, maintenance, and manual coordination exceed the time saved by correlation. You've hit it when correlation works and your incident lifecycle still runs on human glue: Slack channels, Google Docs, and copy-paste between five tools.
Thoughtworks found that across its 2025 AIOps deployments the highest-value use cases were about knowledge, not autonomous actions: detecting duplicate incidents, retrieving operational knowledge, and assisting with root-cause analysis. That's the ceiling. AIOps gives you knowledge. It doesn't give you resolution.
The check: MTTR has held flat for two or more quarters while alert volume has fallen.
This is the clearest signal that correlation isn't your bottleneck. Observability tools give you data, not answers, and the investigation loop eats most of MTTR when context lives across several tools. Engineers run manual scavenger hunts: alert fires, open the first dashboard, note the timestamp, hunt through logs, search code, repeat.
MTTR has multiple phases: detection, acknowledgment, assembly (the coordination phase where you create channels, find experts, and gather context), diagnosis, and resolution. AIOps improves detection. It doesn't touch assembly, diagnosis coordination (who's looking at logs versus metrics), or resolution documentation (what did we try and what worked).
Manual team assembly costs 10 to 15 minutes per incident. Across 15 to 20 incidents a month, that's 150 to 300 minutes of coordination overhead before a single line of code gets touched.
The fix isn't better correlation. It's collapsing the gap between "we know something is wrong" and "the team is working on it." That means auto-creating incident channels when an alert fires, auto-paging the on-call rotation, and starting timeline capture immediately.
Favor reduced MTTR by 37% using incident.io, largely by eliminating manual coordination overhead.
The check: writing the post-mortem takes longer than fixing the incident did.
Someone has to scroll through Slack threads, check alert history, review Datadog dashboards, and rebuild the timeline from memory.
Post-mortem archaeology is the manual, time-consuming process of reconstructing incident timelines after the fact. When your AIOps dashboard shows a correlation graph but can't tell you who decided to restart the service, or when, or why, you're still doing archaeology.
Manual post-incident reconstruction runs 60 to 90 minutes per incident as teams search through chat history, monitoring tools, and call recordings trying to piece together what happened. AIOps doesn't solve this because it stops at detection and correlation. It doesn't capture decisions, role assignments, or customer communications.
If tools don't talk to each other, you're forcing engineers to be the manual API, copy-pasting context between tabs while production burns. AIOps knows an incident happened. It doesn't know what your team did about it.
The critical decisions happen in Slack or on a Zoom call. AIOps doesn't see those. When you write the post-mortem three days later, you're reconstructing not just what the monitoring data showed, but what the team discussed, what fixes were ruled out, and why the incident commander chose option B over option A.
Our Scribe joins the call, transcribes it in real time, and flags key decisions as they happen. The timeline then builds itself from those captures plus every Slack message, /inc command, and role change. We draft the post-mortem from that timeline, roughly 80% complete before anyone opens a document. Engineers spend 15 minutes refining the draft instead of 60 to 90 minutes reconstructing events.
The check: alert volume is down and your rotation is no less stretched than it was.
Better filtering didn't help because the problem was never the alerts themselves.
On-call friction comes from coordination overhead more than technical complexity. When engineers must locate runbooks, identify responders, create communication channels, and update stakeholders manually, those steps delay troubleshooting before investigation even begins. That overhead exists whether you get 500 alerts or 50.
Clunky tooling burns people out: responders spend the first minutes of an incident setting up meetings and channels instead of solving the problem. When every incident starts with "create a channel, ping these four people, open these six tabs," the pager still feels heavy even when it fires less often.
The trend is going the wrong way. The SRE Report 2025 puts median toil at 20% of work, up from 14% the year before and the first rise in five years. Median time on operations activities rose from 25% to 30% over the same period. AIOps reduces alert noise, but it doesn't reduce the coordination overhead that happens after the alert fires.
Burnout drops when the incident process is structured and automatic. Google's incident management guide states that automating elements of incident response will free on-callers to focus on problem solving, including automation of common tasks, automated analysis of key impact information, root cause analysis, and intelligent suggestion of mitigating actions.
That means the incident channel creates itself, the on-call engineer gets paged automatically, the timeline starts recording immediately, and the post-mortem drafts itself. Burnout comes from chaos and uncertainty, not just volume.
The check: you can't answer a leadership question about one service's reliability without a manual export.
Leadership asks why the payments service keeps breaking. You export data from three tools and build a spreadsheet.
AIOps dashboards show alert patterns and correlation graphs. Many of them track MTTR and incident frequency, but the data sits across PagerDuty, Datadog, Slack, and Jira, which makes it hard to answer leadership questions about service reliability trends.
Your AIOps platform knows 500 alerts correlated into 12 incidents. It doesn't know who was involved, what the customer impact was, or which follow-up tasks got completed. That's not what it was built for.
incident.io's Insights dashboard provides MTTR trend and incident-pattern data, giving you a concrete answer when leadership asks if the SRE function is working. The data is captured automatically during the incident, not reconstructed afterward.
The visibility gap is the difference between knowing an incident happened and knowing how your team responded. Any AI-generated incident data needs human review before it becomes the official record.
To maintain accountability and factual integrity, AI systems with knowledge bases should enforce a strict provenance rule: every factual assertion must be linked to an authoritative source. In incident response, that means the timeline is captured in real time from actual Slack messages and status updates, not generated from model inference three days later.
The check: you run custom code whose only job is moving AIOps output into your incident process.
Your team built a homegrown bot that takes AIOps output and posts it to Slack, creates a Jira ticket, and updates a Google Sheet. It worked at first. As incident volume grows, nobody wants to maintain the bot.
The correlation half works. The wiring that connects it to your process is yours to keep alive.
AIOps platforms surface their output in a dashboard. Your team works in Slack. Every incident requires someone to copy the correlation summary from the dashboard, paste it into a Slack channel, manually tag the on-call engineer, and then start a Zoom call. That's not an incident process. That's a series of manual hand-offs.
incident.io runs the incident lifecycle in Slack with /inc commands. No copy-paste, no context switching, no web UI bolted onto chat.
Homegrown incident bots are maintenance debt. The person who built it moved to a different team. The AIOps vendor changed their API. The Slack integration broke after a permissions update. Now your incident process depends on a tool nobody owns. Our native integrations cover the alert-to-post-mortem path, so the glue code you're maintaining today has nothing left to do.
The check: new on-call engineers need weeks of shadowing before you trust them to run point alone.
Your new on-call engineer's first incident goes badly not because they lack technical skill, but because the process lives in tribal knowledge: which Slack channel to create, who to page, which runbook to follow, when to escalate.
AIOps doesn't help here because it doesn't encode your incident process. It just tells you an incident is happening.
AIOps platforms assume you already have a mature incident process. They don't teach new engineers how to run an incident, when to escalate, or how to write a post-mortem. That's out of scope.
Google's SRE Book found that thinking through and recording the best practices ahead of time in a playbook produces roughly a 3x improvement in MTTR compared to the strategy of winging it. The practiced on-call engineer armed with a playbook works much better than the hero jack-of-all-trades engineer.
Tribal knowledge doesn't scale. When your senior engineer goes on vacation, your incident process goes with them. Our /inc commands, structured triage, and Catalog context make the process repeatable, so new engineers run their first incident with guidance.
"Super customizable, including automations... SSuper configurable on-call rotation options that let's us create "training" schedules for new joiners." - Verified user on G2
The check: alert thresholds, correlation rules, or model retraining land on your sprint board every month.
Your AIOps tool needs constant tuning: alert thresholds drift, correlation rules break when you deploy new services, the ML model needs retraining. None of that work shows up as incident response, but it comes out of the same budget of hours.
Alert correlation models drift as your infrastructure changes. New microservices, new dependencies, new alert sources. Each change requires tuning: adjust thresholds, retrain models, update correlation rules. That tuning is real work, and it competes with actual incident response.
Of the 20 AIOps proof-of-concept projects Thoughtworks delivered in 2025, 11 reached production. The rest failed for structural and technical reasons. Continuous tuning was one of them: operations teams lack the capacity to run and improve intelligent systems.
If your team spends more hours troubleshooting the AIOps tool than resolving incidents with it, you've outgrown the category for incident response. The tool has become the toil it was supposed to eliminate.
At that point, tuning harder won't help. The gap isn't in correlation quality. It's in everything the correlation output doesn't do.
When AIOps fails during a P0, your rotation needs a fallback that doesn't depend on it. Agree these three steps before you need them:
Write the protocol down and the next AIOps failure costs you minutes instead of a stalled response.
AIOps and AI SRE address different parts of the incident lifecycle. AIOps focuses on correlation and filtering. AI SRE extends to investigation, coordination, and action.
| Capability | AIOps | AI SRE |
|---|---|---|
| Alert correlation | Yes, via integrated monitoring | Yes, via integrated monitoring |
| Root cause analysis | Surfaces patterns, human interprets | Investigations analyzes telemetry, code changes, past incidents |
| Fix generation | Varies by platform | Drafts fix pull requests for human review |
| Coordination | Basic workflow automation | Auto-creates channels, pages on-call, captures timeline |
| Post-mortem drafting | Some platforms offer AI summaries | Auto-drafts from captured timeline in 15 minutes |
| Incident lifecycle coverage | Detection through learning, varying depth | Detection through post-mortem and follow-up tracking |
The distinction: AIOps tells you something is wrong. AI SRE tells you what's wrong, who's fixing it, what they tried, and drafts the post-mortem.
The incident lifecycle includes multiple phases: detect, acknowledge, assemble, diagnose, resolve, document, and learn. AIOps handles detection well. Everything else requires coordination unless you have an incident management platform.
We're the Slack-native incident management platform that eliminates coordination overhead by auto-creating channels, auto-capturing timelines, and auto-drafting post-mortems. Your AIOps tool still handles detection and correlation. We handle the other phases.
AIOps platforms don't know who your on-call engineer is, which Slack channel to create, or what your post-mortem template looks like. That context lives in your incident management process, not your monitoring stack.
FireHydrant's AI summarizes incidents in real time, captures call transcripts, drafts retrospectives, and suggests triage steps and follow-up actions, on its Enterprise tier. Rootly's AI SRE ingests logs, metrics, traces, and deployment history, then surfaces ranked root-cause hypotheses, with every remediation action gated on engineer sign-off. Our Investigations goes one step further: it connects telemetry, code changes, and incident history to identify root causes, then drafts the fix itself as a pull request a human reviews and merges.
Investigations automates up to 80% of incident response, so engineers spend their time on the fix rather than the search. It is a paid add-on to Response, available on Pro and Enterprise.
The human stays in the loop. Google's SRE autonomy levels put assisted automation at L1: the agent analyzes data and provides insights or suggestions, but a human approves and actuates any action. Investigations never takes action on production systems without that review.
Vanta automated a manual 5-step incident process. Skyscanner replaced fragmented Jira and Slack workflows, so non-technical teams now run incidents themselves. Etsy reported incident.io shipping 4 requested features in the time a competitor answered 1 support ticket. Fin migrated off PagerDuty and Atlassian Status Page in a matter of weeks.
The pattern: AIOps handles the correlation problem it's designed for. incident.io handles the coordination and action problem AIOps was never built to solve.
Your AIOps tool groups alerts. Investigations takes the next three steps: triage, root cause, drafted fix. Your team reviews and merges. Book a demo and watch us take it from alert to drafted pull request.
AIOps: Artificial Intelligence for IT Operations. Platforms that use machine learning to analyze telemetry, correlate events, and surface patterns for IT operations teams, focused on noise reduction and anomaly detection.
AI SRE: a category of tools that automate incident investigation, root cause analysis, and remediation, moving beyond correlation to draft fixes and coordinate response. Our product in this category is Investigations.
Alert correlation: the grouping of related alerts into a single incident based on factors like time and topology, reducing thousands of raw events into a few dozen actionable incidents.
Incident coordination: the process of assembling the right team, assigning roles, capturing decisions, and communicating status during an active incident, often in Slack or Microsoft Teams.
MTTR: mean time to resolution, the average time from incident detection to full resolution, covering detection, acknowledgment, assembly, diagnosis, and resolution.


We spent the last two years building Investigations, our AI SRE. In this post, go behind the scenes of the work done and why build vs. buy is one of the most important decisions engineering teams are making right now.


Today we're launching Investigations: agentic root cause analysis that starts the moment you're paged, figures out what broke and why, and works with your team through to resolution. Here's what we built, what's powering it, and why it took some time to get right.


PagerDuty published a new comparison table about incident.io. Once again, it describes a product we don't recognize. So once again, we're correcting the record, row by row, with receipts.

Ready for modern incident management? Book a call with one of our experts today.
