The most-quoted number in AI voice right now is also the easiest one to fake, and you don't even have to cheat to fake it. You just have to use the default definition.

Containment rate is a division problem: contained sessions over total sessions entering the automated channel. Contained means the call ended without a transfer to a human. Read that definition twice and the flaw becomes obvious — it measures the absence of an escalation, not the presence of a resolution. A caller who got a perfect answer and a caller who gave up in disgust and hung up produce the same log signal. Both are wins. Your dashboard cannot tell them apart, and by default it doesn't try.

This isn't a theoretical complaint. The U.S. Postal Service runs one of the largest IVR deployments in the country — the 1-800-ASK-USPS line handled more than 90 million calls in FY2020 — and its own Office of Inspector General audited the containment methodology in a September 2021 white paper. The finding: every call not transferred to an agent was counted as contained, including calls where the customer hung up without resolution. Customer-ended calls made up the majority of contained calls. The OIG concluded the agency was overstating containment and recommended reporting complete contained calls separately from incomplete abandoned ones.

Now look at what the headline number was doing while that was true. USPS containment climbed from 62% of 72 million-plus calls in FY2018 to 78% between October 2020 and May 2021, against a stated target of 80%. Meanwhile, 26% of IVR survey respondents in FY2020 were very dissatisfied and nearly four in ten said their issue went unresolved inside the system. The metric went up and to the right. The experience went down. Neither chart knew about the other.

The gap between "didn't escalate" and "got helped" is enormous

Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service issues fully resolve in self-service. For issues the customers themselves described as very simple, it was 36%. Hold that against the 70–90% containment figures circulating in vendor marketing. The gap between 14% and 80% is not measurement noise — it is the size of the problem. The same survey named the mechanism: 45% of self-service starters said the company didn't understand what they were trying to do, and 43% couldn't find content relevant to their issue. Both of those people leave without requesting a transfer. Both get logged as contained.

We got a live demonstration this year. Per GAO, IRS phone automation answered 41% of filing-season calls in 2026, up from 34% in 2025 on roughly 28 million calls. Over the same period, average wait time went from three minutes to eight. Containment and customer experience moved in opposite directions inside a single fiscal year, on the same phone system.

CallMiner flagged the incentive structure directly in an April 2026 brief, calling it the Cobra Effect: optimize hard enough for containment and you don't build an agent that helps, you build an agent that traps. Deflection has the inverse failure mode — it rewards agents that get rid of people efficiently. On the same deployment, deflection and containment now routinely differ by 20 to 40 percentage points, which is why 2026 RFPs have started scoring them in separate columns. Deloitte's 2026 cross-industry survey pegs average AI voice containment at 41%, with financial services at 52% and healthcare at 29%. If a vendor won't split deflection from containment for you, that refusal is the answer.

Aggregate containment is the part that lies

The single most useful table in that USPS report isn't the headline — it's the per-intent split. General inquiry contained at 93%. Passports 91%. Post office lookup 90%. Then: package tracking 72%, international tracking 63%, redelivery 62%, hold mail 60%, change of address 56%. That's a 37-point spread across eight common intents, one platform, one reporting window.

The pattern is consistent everywhere intents get broken out. Intents that only relay information contain beautifully. Intents that require the caller to supply a tracking number, an address, or a date contain badly, because every additional turn is another chance to lose the thread. So a 78% aggregate is compatible with 93% on store lookups and 56% on address changes — and the address-change callers are the ones churning. My rule: an aggregate containment number with no intent breakdown behind it isn't a metric, it's a mood.

One more thing worth writing down and versioning: your denominator. Two engineers can compute containment from the same log table and land ten points apart depending on what each one excludes. Containment and escalation only sum to 100% when abandonment is reported as its own line. Blend abandonment into containment and the number stops meaning anything. Any change to the inclusion rule starts a new time series, not a continuation of the old one.

What honest containment reporting requires

Here's the structural argument, and it's the reason I think this is a platform problem rather than a reporting problem: you cannot audit containment from routing telemetry alone. Routing knows the call didn't transfer. Only the content of the call knows whether anything got resolved. If your voice AI vendor and your conversation intelligence vendor and your contact center are three different systems, the contained calls are exactly the population nobody owns — the AI vendor counts them as successes and the contact center never sees them, because by definition a human never touched them.

Dial800 CallCenter reports AI Containment as four fields — AI Calls, Contained, Escalated, and Containment % trended over time — but the number is only defensible because the contained calls carry everything else on the same record: a speaker-labeled transcript, a 0–100 sentiment score, an AI summary, and a nightly auto-QA verdict scored against plain-language scorecard criteria with a quoted line from the call as evidence. That's what turns "contained" from an assertion into a claim you can check. Sort contained calls by ending sentiment and read the bottom decile; that population is your real containment error rate, and it takes about twenty minutes to find.

The handoff mechanics matter for the same reason. When an AI agent escalates, it attaches a context summary to the interaction card, so the human picks up mid-story instead of restarting — which means escalation stops being the failure state the metric implicitly treats it as. Agents can hand a caller back to a voice AI agent from the directory for routine follow-through. And because attribution lives on the same platform, a contained call still carries its campaign, keyword, and web journey, so you can ask the question that actually matters: did the calls our AI contained from this campaign turn into revenue, or did we just spend media budget to route buyers into a dead end politely?

Containment is a fine operational metric. It is a terrible success metric. Run it next to abandonment, escalation, and a resolution measure you can defend with a quote from the call — or accept that you're paying per minute to measure your own silence.