Containment rate is the percentage of conversations that ended without a human. It is easy to compute, easy to explain to a board, and — on its own — actively harmful.
The reason is simple. There is a trivial way to raise containment: make it harder to reach a person. Hide the escalation path, argue with the request, answer adjacent questions confidently. Every one of those moves the number up and the business outcome down, and none of them shows up on a containment chart.
Three numbers that keep it honest
We instrument the same three alongside containment on every deployment.
Repeat contact within a window. If a customer comes back within four hours — any channel — the first conversation did not resolve anything, whatever it was scored as. This is the single most useful counterweight, because the behaviours that inflate containment nearly all produce repeat contacts.
Escalation quality. A one-question survey to the agent after a transferred conversation: did you have what you needed to start? Falling scores mean a read path has broken or the handoff payload has lost a field, usually weeks before anything else surfaces it.
Resolution as the customer sees it. Not “did the conversation end” but “did the thing the customer wanted actually happen”: the claim was filed, the seat was changed, the payment posted. Where the action is in a system you can read, this is measurable without asking anyone.
A deployment with 70% containment and 4% repeat contact is healthier than one with 85% containment and 19% repeat contact. The second one is a queue with extra steps.
Segment before you celebrate
An aggregate containment number hides the only thing worth knowing: which intents are working. Password resets contain at ninety-something percent everywhere and tell you nothing. The interesting movement is always in the middle band — the intents where the assistant sometimes helps and sometimes does not — and that is where the read paths and stopping rules need attention.
We report containment by intent, never as a single number, and we exclude the categories that are deliberately never contained so they do not drag the figure down and create pressure to contain them.
Watch what happens after the handoff
Handle time for transferred conversations is a trap in the other direction. It rises when the assistant successfully contains the easy contacts, because what remains is genuinely harder. Teams that do not expect this conclude the automation is making agents slower.
Compare like with like: handle time for the same intent before and after, not the average across a changing mix.
The uncomfortable one
Ask, occasionally, what share of contained conversations the customer would describe as resolved. Sampling a few hundred transcripts by hand, quarterly, catches things no dashboard does — the polite dead end, the answer that was technically correct, the customer who gave up.
It is slow, it does not automate well, and it has changed our roadmap more than once.
