# How Kaizen Gaming automated incident response across 20 markets

At peak, its platform processes more than 800K requests per minute across more than 1,500 microservices. Keeping the engine running is a 25-person Site Reliability team operating 24/7 across two continents, alongside a broader network of more than 200 engineers on the on-call rotation, against an availability objective of 99.95% and an MTTR target of under 12 minutes. At that scale, every minute of incident coordination matters. By bringing incident response, on-call escalation, customer communications, and post-incident follow-up into one connected platform with [incident.io](http://incident.io/), the Site Reliability team has cut the manual work of a major incident and built a response model designed to scale with the business.

## Why Kaizen Gaming chose incident.io

Kaizen Gaming's incident response used to be entirely manual. Someone had to handle every step, from pulling together a response team to updating customers on the status page. During a major incident, that left most of the engineers on shift stuck on coordination instead of the actual problem.

Kaizen Gaming needed more than a replacement status-page solution. The team wanted a connected platform capable of supporting the full incident lifecycle: declaring and coordinating incidents, paging the right responders, communicating with customers, managing follow-up work, and learning from incident data.

To identify the right platform, Kaizen Gaming ran a structured proof-of-concept across several incident-management solutions. The evaluation focused on end-to-end incident management, deep Slack integration, flexible automation and APIs, on-call functionality, regional and localized status-page communications, actionable reporting and insights, ease of adoption, scalability, and strong vendor support. [incident.io](http://incident.io) stood out by bringing these capabilities together in one integrated platform, while giving the team the flexibility to tailor workflows to Kaizen Gaming's operating model.



## Automating the first minutes of a Sev1

Kaizen Gaming's incident response used to be entirely manual. When a Sev1 was declared, the team coordinated every initial step by hand: they created a dedicated Slack investigation channel, posted updates in other channels to give the wider organization visibility, launched a Google Meet incident bridge, and called the relevant responders one by one to bring the right teams together. Completing the initial response setup took approximately five to eight minutes, and assembling the response team through individual phone calls could take another five to ten. All of it happened during the most critical stage of the incident, before anyone was able to focus on diagnosis or mitigation.

> Before incident.io, the first five minutes of a major incident were spent calling engineers one at a time. Now the right response team is assembled in under two minutes.

As Vaggelis Kosvyras, Incident Manager, puts it: "The hardest part of an incident isn't always the technical problem. It's coordinating the response." Today, much of that coordination happens automatically. When a Sev1 is declared, [incident.io](http://incident.io) creates a dedicated response space and captures the information needed to coordinate the incident. Kaizen Gaming's configured Workflows then create the Slack channel, assign incident roles, notify the relevant teams, page the right people through On-call, and create or update the status pages, all automatically. Using On-call schedules, escalation paths, and Catalog relationships, the incident is routed to the appropriate responders based on the affected services and the assessed impact, within two minutes.

As a result:

* The initial response workflow launches in under two minutes, compared with five to ten minutes previously.
* The required responders are assembled in about two minutes, compared with five to ten minutes of individual phone calls.
* Relevant stakeholders across the company, including senior leadership, country managers, Customer Support, operational teams, and affected business functions, are notified within one to two minutes through automated announcements and targeted Slack notifications.
* Engineers joining an incident already in progress use Scribe's real-time summaries to catch up without interrupting responders.

> Approximately 80% of our initial incident-response workflow is now automated. The team can spend less time administering the incident and more time resolving it. 



## Status pages that update themselves

Kaizen Gaming previously relied on a standalone solution to manage 20 customer-facing status pages across its markets and operators. API limits required updates to be processed in batches, so an incident affecting multiple markets could involve several rounds of manual work. And because the status pages weren't connected to the internal incident record, every material change had to be recorded internally and then communicated separately across all affected pages.

The cost of that gap showed up during a major infrastructure incident. With internal systems and response tooling under significant strain, the same tools the team relied on to update customers were the ones failing, so at the moment they most needed to communicate, they couldn't do it reliably.

That's the gap [incident.io](http://incident.io) closed. Kaizen Gaming configured Catalog to map the relationships among services, components, operators, status pages, on-call teams, and escalation paths. When predefined conditions are met, Workflows create or update the relevant status-page incidents automatically as the internal incident evolves, so one connected workflow updates every affected status page, in the right language for each operator, without responders switching between tools. In an outage like the one that first exposed the problem, customer communications now keep flowing straight off the incident record.

The same connected model has also simplified the creation of new status pages. Previously, launching a page required procurement, billing approval, content preparation, and coordination across Procurement, Finance, Billing, Product, and Engineering, and typically took four to six weeks. Today, the equivalent end-to-end process can be completed within one working day.

> "It used to take us between four and six weeks to launch a new status page. Today, the equivalent process takes roughly one working day. That time goes back to the team so they can focus on higher-priority work."



## Bringing structure to post-incident work

After an incident, follow-up actions used to be captured across Jira tickets, Slack conversations, and Confluence post-mortem pages, with no single, standardized way to track ownership, due dates, reminders, and completion. That made it hard to ensure agreed improvements actually resulted in action.

Now, regulatory, business, technical, and post-incident reporting is centralized in [incident.io](http://incident.io). Follow-up actions have clear owners, priorities, and due dates, and [incident.io](http://incident.io) Policies notify owners as deadlines approach and when tasks become overdue, giving the team centralized visibility over all outstanding work.

> Before, assigning a follow-up action was no guarantee that it would be completed. Now, incident.io automatically notifies the owner, and the follow-up process is as structured as the response itself.

## Turning incident data into operational insight

Using [incident.io](http://incident.io) Insights, Kaizen Gaming gained a clearer view of when incidents occurred, which services were most frequently affected, and how much responder time was spent managing them. When incident data lived across disconnected tools and manual notes, that view didn't exist. "Before, we didn't have a clear picture of our own incident patterns. Now we do," says Vaggelis. Instead of relying on assumptions about when and why incidents happened, the team works from consistent data, and it's helping them move from reactive response toward proactive improvement. That's already shaping how Kaizen Gaming thinks about release timing, maintenance windows, and change management.

## Implementation and adoption

Following a one-month phased rollout, Kaizen Gaming migrated its incident-response workflows and 20 status pages to [incident.io](http://incident.io). The work was completed through a cross-functional effort involving the Site Reliability Operations (SRO), Information Security (IS), and System Administration (SYS) teams, together with Site Reliability Engineers (SRE), Network Reliability Engineers (NRE), and Incident Managers (IM).

Kaizen Gaming chose to complete the status-page migration before the standalone solution's contract expired, despite having several months remaining, a decision that reflected how much value they saw in connecting everything within one platform.

## Looking ahead

For a Site Reliability team supporting a complex, always-on platform, the most important result isn't automation for its own sake. It's the ability to reduce coordination overhead, bring the right people into an incident sooner, communicate more consistently, and give responders more time to focus on mitigation and recovery.

> "Adopting incident.io has significantly improved the way we manage incidents day to day. Faster coordination reduces operational overhead, but the bigger benefit is allowing our people to focus on what matters most."

Kaizen Gaming plans to keep building on this connected model: extending automation across its incident workflows, using Insights to guide operational decisions, and expanding the same approach to new markets and teams as the business grows. As it enters new markets, the incident-management model can scale without every additional operator, language, or regulatory requirement introducing another disconnected process.