AI Agents
Containment Is the Wrong Metric for Self-Service
Containment counts customers who gave up alongside customers you helped. If it is your headline self-service metric, you are optimizing for the wrong outcome.
6 min read
Containment — the share of contacts that finish in self-service without reaching an agent — is the most commonly reported measure of automation success. It is also the easiest to improve for the wrong reasons.
A customer who gets their answer and hangs up satisfied is contained. So is a customer who could not find the option they needed, could not reach a human, and gave up. Both look identical in the metric. One of them will call back tomorrow, or churn.
What to measure instead
No single metric replaces containment, but four together are considerably harder to game.
- Resolution rate — contacts where the customer’s stated intent was actually completed, verified against the system of record where possible.
- Repeat contact within a window — the share of contained contacts that generate another contact from the same customer within 24 or 72 hours.
- Escalation reason — not just how often the agent escalated, but why, categorized well enough to act on.
- Cost per resolved contact — total cost of the automation divided by resolutions, not by contained contacts.
The second of those is the most revealing and the least often instrumented. A high containment rate paired with a high repeat-contact rate is a system that is deflecting rather than resolving, and it is usually cheaper to fix once you can see it.
The escalation path is part of the design
Teams optimizing for containment tend to make escalation harder — burying the option, adding friction, requiring the customer to try automation twice. This raises containment and lowers customer satisfaction, which is a trade a metric can hide and a customer cannot.
A better posture: make escalation easy, and make automation good enough that customers do not want it. Where escalation is easy, containment becomes an honest measure of whether the automation is genuinely helping, because customers who are not being helped will leave.
This gets more important with generative AI, not less
A traditional IVR fails visibly — the customer hits an option that does not fit and presses zero. A generative agent fails plausibly. It produces a fluent, confident, wrong answer, the customer accepts it, and the contact is contained. The failure surfaces later, as a repeat contact, a complaint, or a bad outcome the customer acted on.
Fluency makes failure quieter. Instrumentation has to get louder to compensate.
This is why grounding answers in approved content, maintaining evaluation sets, and tracking repeat contact matter more once generative models enter the path. The failure mode moved from visible frustration to invisible misinformation, and your measurement has to move with it.
More insights
- Amazon ConnectSequencing an Amazon Connect Migration So No Single Step Can Hurt YouMost contact center migrations fail at the cutover, not the build. The fix is sequencing — choosing an order of migration where every phase is small enough to reverse.
- Agentic AIWhat Agentic AI Needs Before It Goes Anywhere Near ProductionBuilding an agent that acts is straightforward. Building one you can operate, audit, and stop is the engineering problem — and it is the part most pilots skip.