How We Cut Meridian Pay’s Incident Rate 80% in 11 Weeks
Meridian Pay, a fintech in Austin, was waking up to PagerDuty alerts every night. Engineers were burnt out, customers were hitting failed payments, and nobody trusted the payment pipeline anymore. They hired us to fix it.
Eleven weeks later, their incident rate was down 80%. This is the honest story of how we got there — and what it means for teams still stuck in the same loop.
The Problem: Constant PagerDuty Alerts
The incident logs told the story. Dozens of alerts a week, most of them the same handful of recurring failures:
- Timeouts on payment processing that spiked at peak load
- A database connection pool that exhausted under traffic bursts
- Noisy alerting that desensitized the team — real problems got lost in the noise
- A deploy process that occasionally pushed broken config to production
The team wasn’t bad at their jobs. They were drowning. Every incident got patched, but the root causes never got fixed, so the same fires kept starting. That’s the classic incident loop: treat symptoms, never rebuild, never sleep.
The Audit: Finding the Fragile Systems
We didn’t start with code. We started with the incident history. We pulled six months of PagerDuty data, grouped alerts by pattern, and ranked them by frequency and business impact.
Three findings stood out:
- The payment pipeline was fragile. The most critical path in the business was also the least resilient. A single upstream timeout cascaded into failed payments for every customer behind it.
- Observability was almost nonexistent. They had monitoring, but it answered “is it down?” not “why is it slow, and for whom?” There was no tracing, so debugging took hours of guesswork.
- Alerting was tuned wrong. Too many low-severity alerts trained everyone to ignore the tool. The important signals were indistinguishable from the noise.
The audit was the easy part. The hard part was getting everyone to agree the rebuild was worth it. Our answer: we didn’t ask them to trust us — we asked them to watch the data.
The Rebuild: Architecture That Scales
We rebuilt the payment pipeline in layers, starting with the highest-risk failure modes and shipping fixes every week so the team saw progress (and fewer alerts) early.
- Retries and timeouts with backoff. We introduced exponential backoff and circuit breakers so a single failing dependency couldn’t take down the whole pipeline.
- Resilient database access. We right-sized the connection pool, added pooling under load, and moved the hottest queries to paths that wouldn’t exhaust connections.
- Deploy safety. We added staging gating and gradual rollouts so broken config couldn’t reach production in a single click.
- Tuned alerting. We killed the noisy alerts, kept the ones that mattered, and set severity based on business impact — not tooling convenience.
Each change shipped with tests and a runbook, so the team could support it themselves.
The Observability: Catching Problems Before Customers Do
This was the change that made everything else possible. We implemented distributed tracing across the payment flow, so when a request was slow, the team could see exactly which service was responsible — in minutes, not hours.
We built dashboards around the metrics that matter: payment success rate, latency percentiles, and error rates by service. And we tied alerting to those metrics, so the team got paged for real problems with real context — not for a single slow request at 3am.
The result wasn’t just fewer incidents. It was a team that finally trusted their tooling. When an alert fired, they knew it meant something.
The Results: 80% Fewer Incidents
Within 11 weeks, Meridian Pay’s incident rate was down 80%. The recurring failures that dominated their incident history were gone. PagerDuty got quieter. Engineers got their nights back.
More importantly, the team kept the improvements. We transferred knowledge as we went, left runbooks behind, and handed over ownership of the new observability stack. They didn’t just get a fix — they got a system they could maintain.
Key Takeaways for Other Teams
If you’re stuck in the same loop, these are the lessons we’d want you to take away:
- Start with the incident history, not the code. Your PagerDuty logs will tell you exactly which systems are fragile. Rank by frequency and impact, then fix the worst first.
- Fix root causes, not symptoms. Every patch that doesn’t address the underlying failure is a future 3am page. Rebuild the fragile parts properly.
- Observability is the multiplier. You can’t reduce incidents you can’t see. Tracing and honest metrics turn debugging from archaeology into a five-minute task.
- Alert on business impact, not tooling noise. If every alert is urgent, none of them are. Tune severity so the team actually trusts the pages.
Want the Same Result?
If your team is tired of waking up to PagerDuty, we’d love to audit your incident history. Book a 30-min fit check and we’ll tell you honestly whether we can cut your incident rate 80% in 11 weeks.