24/7 site reliability engineering for a digital bank
Observability, SLOs and round-the-clock on-call support for a mobile-first bank with millions of customers.

The client
Kuda — digital banking, Lagos, Nigeria & London, UK.
The challenge
A bank cannot have a maintenance window. Kuda's product team was growing quickly, but reliability still depended on developers noticing problems in dashboards and fixing them out of hours. Alerts were noisy, incidents were resolved ad hoc and there was no shared definition of what "healthy" meant for each service.
What we did
We introduced a site reliability practice and staffed it around the clock together with Kuda's engineers.
- Service level objectives for the customer-facing journeys: sign-up, transfers, card payments, statements
- Consolidated observability with Prometheus, Grafana and the Elastic stack; alert rules rewritten around SLO burn rates
- Opsgenie on-call rotations with escalation policies and runbooks for every alert
- Blameless post-incident reviews and a weekly reliability report to leadership
- Backup and cross-region recovery drills every quarter
Results
Customer-facing incidents fell by 62 percent in the first six months, mean time to acknowledge an alert is under five minutes, and the engineering team spends its nights sleeping instead of watching dashboards.
We needed 24/7 coverage without hiring a 24/7 team. Their SRE service gave us real monitoring, real on-call and a lot more sleep.
— Chief Technology Officer, Kuda
More cloud engineering work
A Kubernetes platform built for millions of daily transactions
Platform engineering, GitOps and 24/7 SRE for one of Nigeria's largest payment and business banking platforms.
Read the case study →Cutting a multi-country e-commerce cloud bill by more than a third
FinOps, right-sizing and autoscaling across the platform that serves Jumia's Nigerian marketplace.
Read the case study →Zero-downtime migration of airline booking systems to AWS
Moving reservations, operations and customer-facing web systems from a co-located data centre to AWS over a single weekend.
Read the case study →