Join live: What happens when your AI goes down?
Join live: What happens when your AI goes down?

At incident.io, we've spent the last two years building Investigations, our AI SRE. When you get paged, it starts investigating straight away, looking across your telemetry, recent deploys, past incidents, docs and code, and posts what it's found in your incident channel (or on your phone, if it's 2am and you're still deciding whether you need to get out of bed). By the time you open your laptop, you're starting at step six of triage rather than step one.
I've worked on Investigations since pretty much the beginning, so I'm going to let you in on a secret: diagnosing incidents with AI is a whole lot harder than it looks. What a responder sees is just the tip of the iceberg, and underneath it we've had to build a load of new systems to make sure it's consistently accurate, fast and useful.
In this post I'll take you behind the scenes of what that took. Build vs buy is one of the most important decisions engineering teams are making right now, and I think it's really worth having a clear picture of what building something like this involves before you commit to it.
In November 2024, Lisa, one of our engineers who'd been working on an early version since that August, sent the team a message about what we were pretty sure was our first holy sh*t moment. It had diagnosed a real incident, and we were bullish that we'd have something in customers' hands within a few weeks.
It took us two years to get to something we were confident in. I promise we're a smart team, and we've been working on this nonstop with as many people as we could sensibly put on it, so it really does just take this long. It's been one of the biggest rollercoasters of my engineering career!

With agentic coding, it's so easy to feel like you can build anything in an afternoon, and a lot of the time you can. We build with Claude and Cursor every day, and the way we write software has changed more in the last six months than in the previous decade. But the gap between an impressive prototype and a reliable product that your whole team leans on is massive, and the shiny prototype is a really easy way to accidentally sign yourself up for months of maintenance. I've spoken to plenty of people who expected to spend a few days building their own AI tooling, and six months later they still had a full-time team keeping it running.
The problems we hit fell into four categories, and we hit them roughly in this order:
There are countless paper cuts in each of these, so rather than listing them all, I'll go deep on one example from each.
With some AI products, it's fine to shrug and say the model is a bit weird sometimes. Incidents aren't one of them. If we send a responder down the wrong path for 30 minutes during a major outage, that's a huge cost to the business in customer trust and engineering time, so we need to know that the system is consistently correct, and that every change we make is moving it in the right direction.
That raises some tricky questions. When you ship a change, how do you know you haven't caused a regression somewhere else? Can you measure the impact of upgrading to a new model? How do you know you're not just tuning to the handful of incidents you happen to be testing on? When we speak to teams who've built their own tooling, one of the first things we ask is how they measure accuracy, and the answer is usually that they don't.
Our answer is backtesting. At its simplest, an incident goes into what we affectionately call our brain, and a hypothesis comes out. We grade every hypothesis against what responders did in the real incident, on a scale where 100% is a bullseye, 65% is one hop short of the real answer, 35% is a miss and 0% is a completely different universe. When we want to make a change, like swapping in a new model for a single prompt, we freeze past incidents in time and run the new version of the brain against exactly the same inputs, so it's a completely fair test. We do this across a whole set of incidents rather than just one, because incidents vary so much that a single good result doesn't tell you very much.

The really difficult bit is the freezing. There's an incredibly fine line between the context the brain needs to do a fair job and context that's quietly giving it the answer. It needs everything a responder had when the incident started: the alert, the telemetry, the PR that caused it, and the past incidents that might hold clues. But the future leaks in everywhere. The fix PR has since been merged into your codebase, someone has written up the post-mortem in Notion, and a responder posted the root cause in the channel. Include any of these and you'll think you're doing brilliantly when really you're just cheating.

Fixing all of this has felt a bit like a video game, where every level is a new way to fool yourself. Sometimes you flatter yourself by reading a PR from the future. Sometimes you do badly because traces have fallen out of the retention period and you never had the right information in the first place. And sometimes you score an incident badly, dig in, and find that it was the humans who were wrong! It's slow work, but it means that when we ship a change, we can say with confidence that we haven't made things worse. You can't wait for your next critical incident to find out you've accidentally tanked performance.
Once you can measure accuracy, you have a load of data on where you're not doing well. The next problem is being able to do something about it.
It's surprisingly easy to build a black box you can't adapt. You can drop the latest model into an incident, give it lots of context, and be impressed by what comes out. But to build something that's consistently accurate and gets better over time, you need a lot more control. When it sounds 100% confident, should it be, and how would you know? When it picks the wrong answer, can you trace why, and fix whatever it's overfitting to? If it's slow, is it easy to speed up?
Getting this right comes down to the harness, which is the word for how you structure these systems (ours involves hundreds of prompts and some long-running processes). As a very rough overview, when an investigation kicks off we fan out in parallel across every source a human might use to diagnose an incident, like deploys, telemetry and Slack messages. We pull what we learn into findings, synthesize those into a hypothesis, and then an adversarial agent tries to disprove that hypothesis to make it stronger, before we loop back round to the top.

Findings are a good example of how much this changes over time. They're how we squash down hundreds of log lines, metrics, Slack messages and docs into specific things we've learned, and the data model behind them has been through five versions:
We always start with the simplest thing we think will work, and then change it as we find the places it's less accurate or hallucinating. This structure is why every claim Investigations makes in your channel links back to its evidence, and why it's comfortable telling you when it isn't sure.
So now you know whether you're accurate, and you can fix the places you're not. Your next big problem is that people will ignore it anyway!
There's AI-generated text everywhere at the moment, and engineers are expected to read more of it than ever, so it's really easy to just not bother with a new source of information. If a tool is noisy, verbose or badly timed early on, people will tune it out, and once they've decided something isn't worth reading it's incredibly hard to win them back. Even if you're useful 80% of the time, the other 20% can do a lot of damage to trust.
A good example of this is our heads-up messages. The hypothesis is our running summary of an incident, but responders aren't going to keep scrolling back up to check whether it's changed, so when we learn something important, we tell them in the channel. Getting these to be useful has meant building up a whole series of gates that a message has to pass before we speak. None of them are rocket science on their own, but they're the kind of thing you only discover one annoying message at a time.



First come the mechanical checks. Is the incident still open? Is someone else about to send the same thing? Do we have enough conviction in what we're saying? Then come what we call vibe checks, where a message might make sense mechanically but the responders are already talking about it, or the situation is already handled, or we'd just be adding decoration rather than signal. Finally, there's tone tuning: speaking with a level of confidence that matches the evidence, keeping it short enough for a busy channel, making it clear what's changed since we last spoke, and backing it up with links and graphs so people can dig further if they want to.
When it works, it's one of my favorite things about the product. We once had an incident where someone on our free plan was using us to send a load of expensive SMS messages to Burundi. We disabled the organization and started winding down, and then Investigations posted a heads-up: another organization had just signed up and was doing exactly the same thing. That's the difference between putting an incident down thinking everything's fine and catching the next problem within minutes.
Just like hypotheses, we grade every heads-up message, using an LLM to score it for novelty (is it telling people something new?) and accuracy (is it correct?). Every time we upgrade a model or change a prompt, that grading tells us whether we've started getting more annoying.

Say you've got here, with a product that's accurate, useful and concise. Now every other team wants it, and they want it to keep working as their software changes. Incidents run on a huge amount of context, and that context is changing all the time, especially now that everyone is shipping faster with AI. If you rename a service, how long is it before your agents find out? And something that works nicely in a demo might fall over when it meets thousands of repos and an observability stack that looks different for every team.
Telemetry is one of the most interesting examples of this. If you've ever looked at a raw Prometheus response, you'll know it's pretty hard to parse as a human, and we assume our agents aren't so different. So whenever we run a telemetry query, we summarize it for the agent: what the question was, why we asked it, and the key stats like the min and max. We also draw a graph, because just like us, an LLM can spot an interesting spike in a picture far more easily than in an array of numbers.


Then, when you roll it out to production, you meet a whole host of people who are unhappy with you, in order. First it's finance, who've noticed you're burning through tokens and have no control over how many queries you're sending or how big they are. Then the observability team gets in touch because you're flooding Grafana with requests and engineers can't run their own. Then engineers tell you that a query was useful, but without a link or a picture they don't know whether to trust it. And finally security want to know how you're storing all the logs your agent is reading. They're all very sensible concerns, and none of them are things you can sort out in an afternoon.
We've also had to put a lot of thought into memory. The simple version is that when a telemetry query works, the agent stores it as a memory and recalls it in future incidents. The problem is that organizations change, so a memory that was correct once can start lying to you as soon as a service is renamed or a log format changes. So we curate. Every night, a process we call Explore looks at recent incidents and re-runs queries against them to check that our guidance and memories still hold. It rewrites the ones that have drifted, and intentionally forgets the ones that are wrong. When you first connect a data source, we inspect its labels and fields, use your team's existing dashboards to bootstrap the questions you tend to ask, and write ourselves a guidance document. You don't need to remember to tell us when something changes, because we'll pick it up the next day.

All of this sits on top of Nexus, our living model of your production environment, built from your incidents, your systems and your team. We had to build it before Investigations could work the way we wanted.
Anyone can call a model these days. For us, being an AI company means doing all the work around the model: testing our AI with AI, building an agent whose whole job is to argue with our other agent, building systems that keep their own knowledge up to date, and making sure nothing reaches a responder unless we can measure that it's better than what came before.
We're holding changes our customers make to the same bar. If you add a custom skill to Investigations, we'll simulate how it would have performed before you roll it out, and then grade every real use afterwards, so you can see on a dashboard whether it's helping or hurting.
This has been a really brief tour, and each section here is just one deep dive into a single problem out of many. Knowing whether it's right, being able to tune it, making it something people want to read, and keeping it working as your organization changes are all things that sit under the surface of that exciting prototype we saw back in November 2024.
I'm not here to tell you what to do. You know your team and your constraints best, and there are good reasons to build things yourself. But if you're weighing up whether to build your own AI SRE, I'd really encourage you to go in with the whole iceberg in view. For us, it's taken a dedicated team two years, and we're still going.
If you'd like to see where we've got to, you can watch our recent webinar here, or take a look at Investigations in action.


Today we're launching Investigations: agentic root cause analysis that starts the moment you're paged, figures out what broke and why, and works with your team through to resolution. Here's what we built, what's powering it, and why it took some time to get right.


PagerDuty published a new comparison table about incident.io. Once again, it describes a product we don't recognize. So once again, we're correcting the record, row by row, with receipts.


Today, we're launching the Opsgenie Rescue Program to make that landing soft: simplified migration and free overlap so you never pay two vendors at once.

Ready for modern incident management? Book a call with one of our experts today.
