TrueLayer

How TrueLayer made reliability a roadmap priority

TrueLayer's core banking team found itself in a position many engineering organizations might recognize. Their payout product was growing quickly, transaction volumes were rising, and edge cases that had once been rare were becoming more frequent. Engineers were increasingly spending more time responding to incidents rather than shipping product.

Industry
FinTech
Company size
Mid-Market
Use case
Better learning from incidents
Products used
On-call, Response, Status Pages, Catalog

The team could see operational work eating into engineering efficiency, but they had no way to measure how much.

Stefano Iasi, Head of Engineering, wanted data that could show exactly where engineering time was going and help make the case for investing in reliability alongside new product work. He found it in the tool TrueLayer was already using to manage incidents.

When growth became operational debt

Stefano joined TrueLayer around three years ago as a Senior Engineering Manager for the core banking team, responsible for the ledger that tracks every transaction and payment, along with the regulatory obligations that come with moving money.

When TrueLayer launched its Payout product, adoption exceeded their expectations. Some early design decisions had assumed certain edge cases would happen infrequently, but with volume growing exponentially, those edge cases began to appear more often.

The burden on on-call grew, recurring incidents piled up, and frustration ticked up. But when Stefano tried to quantify the problem, there wasn't any data to point to.

Engineers flagged the problem, and rightly so. But there was no data. My message was simple: let's track this, because data is what creates change.
Stefano Iasi, Head of Engineering, TrueLayer

That lack of visibility made planning conversations hard. Engineering argued operational work was consuming too much time. Product had roadmap commitments to deliver. Without data, neither side had an objective way to weigh those tradeoffs.

Joel Oughton, Senior Software Engineering Manager responsible for the ledger and banking partner integrations, joined during that period. As he puts it, there was "a lot of firefighting, a lot of incidents." The team had become good at responding to incidents, but they needed to identify the underlying issues that triggered them in the first place.

Turning incidents into data

TrueLayer was already using incident.io to manage incidents. Rather than introducing another system, Stefano wanted to capture more information where engineers already worked. "It made sense to capture that data there," he says. "It was the most natural place to connect root causes to the incidents themselves."

Around the same time, the team also migrated on-call into incident.io. Previously, responsibilities were spread across PagerDuty, Slack workflows, post-mortem documents, and RCA write-ups.

"It was PagerDuty plus Slack flows, quite a few different places that would capture information. We got rid of PagerDuty and brought it into incident.io. Now the post-mortem docs, the calendar integration, it's all in one place. It's easy to look up an incident six months later and find out everything you need to know: what the cause was, whether the process was followed, which follow-ups haven't been addressed."
Joel Oughton, Sr. Software Engineering Manager, TrueLayer

The process change itself was simple. Engineers completed a handful of custom fields when closing an incident, recording trend classifications, root-cause categories, the provider involved, and how the incident had been detected.

Most important was tracking whether the incident was caused by TrueLayer's systems, an external banking partner, or both.

For Ale Caprarelli, one of the founding members of TrueLayer's cross-team incident working group, the goal was to move beyond individual incidents. The team was "already good at resolving incidents and following the practices in incident.io," Ale says. "The natural next step was to look back over the past month, and the months before, for trends, so we weren't just reacting to the number of incidents."

The team exports incident.io data into a Google Sheet that automatically produces trend reports and charts. That automation, Ale says, "takes the manual effort out, so we can actually spend the time on the analysis."

Those monthly reviews gave engineering something concrete to bring into planning conversations.

The conversations with product went from gut feelings to evidence-based prioritization. Problems that had been ignored for multiple quarters got prioritized because of the root-cause data.
Stefano Iasi, Head of Engineering, TrueLayer

The data also helped the team distinguish between issues they could fix themselves and problems driven by external banking partners. First impressions weren't always right, Joel says. Some incidents that looked like internal issues turned out to be partner issues, and the other way around. Digging into root causes helped the team work out which issues needed a closer look.

The data made conversations with partners more productive. Instead of relying on anecdotes, the team could bring incident trends into QBRs and use them to drive improvements and challenge service levels where appropriate.

Data that changed the roadmap

Once the data started shaping planning decisions, engineering focused on the patterns creating the most operational work.

Many of the recurring incidents weren't simple bugs. They were product gaps that had become visible as transaction volumes increased. Instead of repeatedly responding to those same failures, the team expanded the product to handle those scenarios automatically. Between 2025 and 2026, that work reduced incidents caused by product edge cases by roughly 70%.

We got to a point with a lot less noise, because the bigger trends were all addressed. It became sustainable instead of constant firefighting, and now if something new crops up, we spot it and fix it fast.
Joel Oughton, Senior Software Engineering Manager, TrueLayer

The data changed how work was prioritized.

For example, one bug kept getting pushed back until the data showed it was costing roughly 20 engineer-hours a month. That made it, as Ale puts it, "an easy win to prioritize" when it might otherwise have been difficult to justify.

More broadly, on-call well-being became another factor in planning. Instead of evaluating work only by its implementation cost, the team also considered how many incidents a fix would eliminate and how much operational burden it would remove.

For Stefano, the biggest change was being able to communicate reliability work in business terms.

"It changed what I could communicate upward, to the CTO, the CEO, the commercial organization. The message wasn't that we couldn't ship features. It was that, given what the data was showing, investing in reliability was the right call."
Stefano Iasi, Head of Engineering, TrueLayer

Today, every engineering team tracks incident counts and outstanding follow-ups in monthly operational reviews, while severe incidents are reported all the way to the board.

The workflow has become more connected, too. Incident follow-ups link directly to Jira, making it easy to see what remains open, and teams intentionally reserve capacity to address recurring issues before they become larger problems.

Stefano's advice to other engineering leaders is to quantify the problem without creating more work for the team.

"Visualize the problem and quantify it," he says. "To do that, you need support and automation, it's critical that this kind of effort doesn't drain energy from a team that's already stretched. Having a tool you can customize, extract data from easily, and adapt to your specific needs makes all the difference."


You may also be interested in


The shows you bingemade more reliable with incident.io

Signup image

We’d love to talk to you about

  • All-in-one incident management
  • Our unmatched speed of deployment
  • Why we’re loved by users and easily adopted
  • How we work for the whole organization