Skip to content
PointCaaSCX • CCaaS • Agentic AI

AI Agents

Containment Is the Wrong Metric for Self-Service

Containment counts customers who gave up alongside customers you helped. If it is your headline self-service metric, you are optimizing for the wrong outcome.

6 min read

Containment — the share of contacts that finish in self-service without reaching an agent — is the most commonly reported measure of automation success. It is also the easiest to improve for the wrong reasons.

A customer who gets their answer and hangs up satisfied is contained. So is a customer who could not find the option they needed, could not reach a human, and gave up. Both look identical in the metric. One of them will call back tomorrow, or churn.

What to measure instead

No single metric replaces containment, but four together are considerably harder to game.

  • Resolution rate — contacts where the customer’s stated intent was actually completed, verified against the system of record where possible.
  • Repeat contact within a window — the share of contained contacts that generate another contact from the same customer within 24 or 72 hours.
  • Escalation reason — not just how often the agent escalated, but why, categorized well enough to act on.
  • Cost per resolved contact — total cost of the automation divided by resolutions, not by contained contacts.

The second of those is the most revealing and the least often instrumented. A high containment rate paired with a high repeat-contact rate is a system that is deflecting rather than resolving, and it is usually cheaper to fix once you can see it.

The escalation path is part of the design

Teams optimizing for containment tend to make escalation harder — burying the option, adding friction, requiring the customer to try automation twice. This raises containment and lowers customer satisfaction, which is a trade a metric can hide and a customer cannot.

A better posture: make escalation easy, and make automation good enough that customers do not want it. Where escalation is easy, containment becomes an honest measure of whether the automation is genuinely helping, because customers who are not being helped will leave.

This gets more important with generative AI, not less

A traditional IVR fails visibly — the customer hits an option that does not fit and presses zero. A generative agent fails plausibly. It produces a fluent, confident, wrong answer, the customer accepts it, and the contact is contained. The failure surfaces later, as a repeat contact, a complaint, or a bad outcome the customer acted on.

Fluency makes failure quieter. Instrumentation has to get louder to compensate.

This is why grounding answers in approved content, maintaining evaluation sets, and tracking repeat contact matter more once generative models enter the path. The failure mode moved from visible frustration to invisible misinformation, and your measurement has to move with it.

Have a problem like this?

We would rather talk about your specifics than write in generalities.