Every telecom engineer knows the failover checklist. Two trunks, not one. SIP OPTIONS pings every thirty seconds so dead peers get pulled from the pool before a real call touches them. Retry on timeouts and 5xx, never on a 486 or a 603, because "busy" and "declined" are answers, not failures. Spread your points of presence across regions, because two trunks terminating in the same data center are not redundant — they're one trunk with extra billing.
That checklist is correct. I'd sign off on it. But it optimizes for exactly one metric: did the call complete? And in 2026, that's a dangerously incomplete definition of success.
Here's the failure mode nobody puts on the runbook. Your primary trunk starts returning 503s at 9:14 AM. Health-checking works perfectly. Traffic shifts to the secondary carrier in under a second. Call completion never dips below 99%. Nobody pages anyone. Nobody writes an incident report. And for the next four hours, every call that traverses the backup path arrives with a different signaling profile — a different source trunk, sometimes a rewritten P-Asserted-Identity, occasionally a normalized caller ID where the primary passed it through raw. The calls connected. The records are garbage.
The Blast Radius Nobody Measures
The industry data on outages is genuinely sobering, and it's getting worse in a specific direction. Cisco ThousandEyes tracked between 199 and 386 global network outage events per week through Q1 2026, with a 62% single-week spike at the end of February pushing the total to 386. Tier 1 carriers are not exempt — Arelion took a 1-hour-38-minute hit spanning 18+ countries in March, Lumen oscillated across three U.S. metros in January, and Cogent posted repeat failures centered on the same Denver nodes twice in five weeks.
The cost figures get quoted constantly: EMA Research puts unplanned downtime at roughly $14,000 per minute for midsize businesses and $23,750 per minute for large enterprises, and ITIC found 41% of enterprises reporting hourly costs above $1 million. Those numbers are why you built failover in the first place.
But notice what they measure. They measure unavailability. There is no corresponding industry statistic for "the calls went through but the attribution data was silently corrupted for four hours," because almost nobody instruments for it. It doesn't trip an alarm. It shows up six weeks later as a media buyer wondering why a campaign's cost-per-lead spiked on a Tuesday in August.
Why Failover Corrupts Records Specifically
This isn't hand-waving. There are concrete mechanisms, and they're all consequences of doing failover correctly.
Path-dependent identity. When a call egresses a different carrier, the SIP headers it arrives with often differ. Some carriers pass Caller ID transparently; others normalize, truncate, or substitute. If your analytics layer keys attribution on ANI, a backup path that presents numbers in a different format produces records that don't join against your existing data. Same caller, different key, orphaned record.
Attestation downgrade. STIR/SHAKEN attestation is a property of the originating path, not the call. A call that would have arrived with A-level attestation via your primary may land with B or C via the backup, because the backup carrier doesn't have the same relationship with the originating network. If you're filtering, scoring, or routing on attestation level — and increasingly you should be — your failover just silently changed the input to that logic.
Duplicate delivery from over-eager retry. The classic mistake is retrying on codes that represent real outcomes. Retry a 486 and you may deliver the same call to a second target. Both legs generate CDRs. Now your call volume is inflated, your conversion rate is deflated by exactly the same factor, and your cost-per-call math is wrong in a direction that looks like underperformance.
Timing skew. Simultaneous ring is a legitimate strategy when speed-to-answer matters more than channel efficiency — fork to every target, first to answer wins, the rest get a CANCEL. It's also a machine for generating CDR noise. Every forked leg is a record. Unless your analytics layer understands that four records represent one call, your dashboard shows four.
The Fix Is Architectural, Not Configurational
Here's my actual opinion, and it's the reason I keep coming back to this: resilience and observability are the same engineering problem, and treating them as separate layers is the root cause.
The standard enterprise stack treats them as separate purchases. Carrier and SBC handle resilience. A call tracking vendor bolts on analytics downstream by consuming CDRs and recordings. That seam is precisely where failover events get lost — the resilience layer doesn't know it's supposed to tell anyone, and the analytics layer has no visibility into why a record looks different, only that it does.
The architectural answer is that the routing layer and the measurement layer need to share state. When AccuRoute makes a routing decision — a geographic hop, an overflow cascade, a failover to a backup target — that decision should be an annotation on the call record, not an invisible side effect. The call didn't just happen; it happened this way, for this reason. That's the difference between a CDR and a diagnosis.
This is also where the single-platform argument stops being a marketing line and starts being an engineering one. Dial800 runs the numbers, the routing, and the analytics as one system. Attribution data, VoiceInsights AI transcription and sentiment, and AI Tagging all sit on the same platform that made the routing decision. When a path changes, the record knows. That's not achievable when your SBC vendor and your analytics vendor exchange nothing but a nightly CDR dump.
Three Things Worth Doing This Week
Regardless of platform, this is testable:
- Fail over on purpose, during business hours, and then audit the data — not the uptime. Disable a target, let traffic drain to the backup, and compare a hundred call records from each path field by field. If they don't match, you have a problem you didn't know about.
- Alarm on record shape, not just call completion. A sudden change in the rate of records missing a campaign ID, a GCLID, or a session match is a failover signature. Treat it like a page-worthy anomaly.
- Scope your retries to timeouts and 5xx, strictly. If you're retrying a 480 because "it seemed transient," measure your duplicate rate before you defend it.
The uncomfortable truth is that a well-engineered failover is a silent event by design, and silence is exactly what makes it dangerous to a measurement system. You built the redundancy so that nobody would notice. Congratulations — nobody did, including your analytics.
The right question isn't "did the call survive the outage?" It's "did the record survive the outage?" Those are different systems, and only one of them is on your dashboard.