Evalgent
Back to Blog
Voice AI Evaluation

Task Success Rate vs Containment Rate

Deepesh Jayal
12 min read
Task Success Rate vs Containment Rate

# Task Success Rate vs Containment Rate

Quick answer

Task success rate is the share of calls where the caller's goal got verifiably done. Containment rate is the share of calls handled without a human. They diverge when an agent keeps the call but never solves the problem. Task success is the honest outcome metric; containment is a cost proxy.

Most voice agent dashboards lead with containment. It is the number that goes into board decks and vendor pitches. It is easy to compute and it usually looks good. It also hides the failure that matters most to your customers.

A caller can stay on the line for four minutes, hear a confident answer, and hang up with nothing fixed. Containment scores that call as a win. The agent never transferred, so it counts. Task success scores the same call as a loss. The refund was never issued. Same call, opposite verdict.

This post explains the difference between the two metrics, shows a worked example where they diverge, and gives you a repeatable way to catch "contained but failed" calls. It is closely related to two sibling posts. If you want the full definition of each metric on its own, read our task success rate guide and our containment rate guide. Evalgent measures both on your real calls as the independent evaluator.

What each metric actually measures

The two metrics answer different questions. One is about the caller. The other is about your staffing.

Task success rate: the percentage of calls where the caller's goal was verifiably completed, checked against evidence rather than the agent's own claim. Containment rate: the percentage of calls the agent handled to the end without transferring to a human.

Task success is an outcome metric. It looks at what the caller wanted and asks whether they got it. The refund is issued or it is not. The appointment is booked or it is not. The concept traces back to task analysis, which breaks work into concrete, checkable goals. You score the goal, not the conversation around it.

Containment is a cost-and-efficiency metric. It looks at who did the work. If the agent finished the call alone, the call is contained. If a human took over, it is not. Containment tells you how much labor the agent removed from your queue. It says nothing about whether the caller left happy.

Both are legitimate key performance indicators. The mistake is treating them as the same indicator. They measure different things and they can move in opposite directions.

Why teams confuse the two

The confusion is natural. On a good day the two metrics agree. The agent handles the call alone and solves the problem. Contained and successful. Everyone is happy and nobody checks the difference.

The trouble starts when you optimize. Containment is trivial to measure from call logs. Task success needs someone to read the transcript, check the system of record, and judge the outcome. So teams reach for the easy number and quietly let it stand in for the hard one.

That substitution is where Goodhart's law bites. When a measure becomes a target, it stops being a good measure. Push containment as the goal and the agent learns to keep callers on the line. It stops offering transfers. It answers everything, well or badly. Containment climbs. Task success can fall at the same time.

A proxy metric) stands in for something you cannot measure as easily. Containment is a proxy for good service. It is a useful proxy only while it tracks the real thing. The moment you reward the proxy directly, the link breaks. This is the core reason to measure both and never let one hide the other.

How a high containment rate hides a low task success rate

Picture an agent with a 90 percent containment rate. On paper it automates nine calls in ten. Leadership is pleased. Now read the transcripts.

Some of those contained calls ended with a solved problem. Good. Others ended with the caller giving up. The agent looped, misunderstood, or gave a confident wrong answer. The caller hung up rather than fight it. That hang-up is still contained, because no human was involved. But the goal was never met.

This is the "contained but failed" call. It is invisible to a containment-only dashboard. The call never escalated, so it never raised a flag. The caller absorbed the failure silently, and many of them will call back, or worse, churn. Poor handling like this is exactly what erodes customer service trust over time.

The gap between the two numbers is the story. A 90 percent containment rate paired with a 60 percent task success rate means roughly a third of your automated calls are quiet failures. The bigger the gap, the more the containment number is flattering you. Tracking task success alongside containment is the only way to see the gap at all.

A worked example where the two diverge

Consider 100 billing calls to a voice agent. The numbers below are illustrative, chosen to show the mechanic clearly, not measured from a real deployment.

Eighty calls never touch a human. Of those 80 contained calls, 55 end with the billing issue actually resolved. The other 25 end with the caller stuck or hanging up. Twenty calls transfer to an agent. Of those, 15 get resolved by the human and 5 are abandoned in the queue.

Containment rate is 80 out of 100, or 80 percent. That is the headline number. Task success rate is the resolved calls over the total: 55 contained-and-resolved plus 15 escalated-and-resolved, so 70 out of 100, or 70 percent.

Now look closer. Of the 80 contained calls, only 55 succeeded. That is a 69 percent success rate inside the calls the agent kept. Twenty-five contained calls, a quarter of the "wins," were failures the dashboard never surfaced. Meanwhile the transfers, which containment counts against the agent, resolved 15 of 20 problems. The channel that looks worst on containment did honest work.

If you had optimized for containment alone, you would push those 20 transfers down. Some of the 15 resolved-by-human calls would become contained-but-failed calls. Containment would rise toward 90 percent. Task success would fall. You would be automating failure and reporting it as progress.

Four call scenarios, scored two ways

Every call falls into one of four buckets. The table below shows how each scenario scores against the two metrics. This is the pattern to internalize: contained and successful are not the same column.

Call scenarioTask successContainmentWhat it tells you
Resolved and containedSuccessContainedThe ideal call. Agent solved it alone.
Contained but failedFailureContainedThe hidden problem. Agent kept the call, caller left unhelped.
Escalated then resolvedSuccessNot containedHonest work. Human fixed it after the agent handed off.
AbandonedFailureContainedCaller gave up. Counts as contained, helps nobody.

Two rows deserve attention. "Contained but failed" is the row that flatters containment and damns task success. "Abandoned" is worse, because a hang-up counts as contained by most definitions even though the caller quit in frustration. Both rows are invisible if you watch containment alone. Our containment versus deflection guide covers how these edge cases get miscounted in practice.

Why task success is the honest outcome metric

Task success is stricter than containment on purpose. It requires proof. You cannot score a call as a success because the agent sounded confident or the call was short. You check the outcome against evidence.

That evidence is the transcript, the audio, and your system of record. If the caller asked for a refund, the refund shows up in the billing system or it does not. The agent's closing line, "your refund has been processed," is a claim, not proof. Task success verifies the claim. This is the same discipline dialogue researchers used decades ago, when the PARADISE framework tied dialogue quality to whether the user's task was completed, as Walker and colleagues described in 1997.

Because task success measures the outcome, it is hard to game. You cannot inflate it by keeping callers on the line. You can only raise it by solving more problems. That is why we treat it as the north-star metric in our voice agent metrics scorecard. Everything else on the dashboard explains why task success moved.

Task success is also the metric closest to churn and repeat-call rate. A failed task is a caller who did not get helped. Some fraction of them call back, escalate through another channel, or leave. Containment cannot predict that. Task success can.

Why containment is still worth tracking

None of this makes containment useless. It is a real efficiency signal and you should keep it. Containment tells you how much call volume the agent removed from human queues. That maps directly to cost.

The right frame is complementary, not competitive. Containment measures cost. Task success measures value. You want a high containment rate on the calls the agent actually solves, and honest escalation on the calls it cannot. An agent that contains 95 percent and solves 60 percent is cheaper and worse than one that contains 75 percent and solves 85 percent.

Containment also flags the reverse failure. A very low containment rate means the agent transfers too eagerly and the automation is not paying off. So you read the two together. Containment sets the cost ceiling. Task success tells you whether the cost bought anything. For the closely related distinction between containment and resolution, see our containment versus resolution post, which draws that line in detail.

How to use both metrics together to find contained failures

The point of tracking both is to surface the "contained but failed" bucket and act on it. Here is a repeatable process.

1. Define success per call type first. Write a binary, checkable criterion for each reason a caller phones in. "Refund issued in the billing system" is checkable. "Caller satisfied" is not. Do this before you measure anything.

2. Compute containment from your call logs. Mark each call as contained or transferred. This is the cheap number and you likely already have it.

3. Sample the contained calls and label task success. Pull a representative sample of contained calls. For each one, check the transcript, the audio, and your system of record against the success criterion. Score success or failure on evidence, not the agent's claim.

4. Isolate the contained-but-failed subset. Filter for calls that were contained and scored as failures. This is the hidden problem set. Read every one of them.

5. Cluster the failures by cause. Group them: wrong information, unhandled intent, endless loops, or callers abandoning. The clusters point at what to fix in prompts, tools, or escalation rules.

6. Set an escalation floor, not a containment ceiling. Require the agent to offer a human when confidence is low. Accept that this lowers containment slightly. It converts silent failures into honest transfers.

7. Re-measure and watch the gap. Track containment and task success together on every reporting cycle. A widening gap is your early warning that the agent is keeping calls it cannot solve.

Run this loop continuously, not once. The gap between the two metrics is the number to watch, and it drifts as prompts and call mix change.

Who should measure this

There is a conflict of interest baked into these numbers. Containment is the metric a vendor most wants to report, because it looks good and it is easy to compute. Task success is the metric that can make a deployment look worse. A party paid to hit a containment target has little reason to surface contained-but-failed calls.

That is the case for an independent evaluator. A third party scores task success against your evidence, with no incentive to inflate the number. It reads the contained calls the vendor would rather leave unread. Our independent voice AI evaluation post explains why separation of the builder and the grader matters, and our voice agent evaluation guide shows where both metrics sit in a full program.

Evalgent does this labeling on your real calls. We measure containment and task success side by side, surface the contained-but-failed subset, and report the gap between the two. You get the honest outcome number, not the flattering proxy.

Frequently asked questions

What is the difference between task success rate and containment rate for voice agents?

Task success rate measures whether the caller's goal was verifiably completed, checked against evidence. Containment rate measures whether the agent handled the call without a human transfer. Task success is an outcome metric about the caller. Containment is a cost metric about your staffing. They can move in opposite directions.

Can containment rate be high while task success is low?

Yes, and it is common. An agent that keeps callers on the line without solving their problems will show high containment and low task success. The gap between the two is the volume of "contained but failed" calls, where no human was involved but the caller still left unhelped.

Is a contained call the same as a successful call?

No. A contained call only means no human took over. A successful call means the caller's goal got done. A call can be contained and failed at the same time, when the agent gives a wrong answer or the caller hangs up in frustration. Only task success confirms the outcome.

Does containment count a hang-up as a win?

Under most definitions, yes. A caller who abandons the call without a transfer is still counted as contained, because no human handled it. That is a known weakness of containment. Task success scores the same abandoned call as a failure, which is why you need both metrics to read the call honestly.

Why is containment called a proxy metric and not an outcome metric?

Containment stands in for good service because it is easy to compute. It is a proxy, not the real thing. It tracks value only while the agent solves the calls it keeps. Reward containment directly and the agent learns to hold calls it cannot solve, breaking the link between the proxy and the outcome.

Which metric matters more, containment or task success?

Task success matters more, because it measures whether callers got helped. Containment measures cost, which is real but secondary. The right approach optimizes for task success and reads containment alongside it as an efficiency check. An agent that contains everything and solves little is worse than one that transfers honestly.

How do you find contained-but-failed calls?

Compute containment from call logs, then sample the contained calls and label task success against the transcript and your system of record. Filter for calls that were contained but scored as failures. That subset is the hidden problem. Cluster those calls by cause to see what to fix in prompts, tools, or escalation.

How is task success rate different from resolution rate?

Task success rate verifies the caller's specific goal against evidence, so it is the stricter, outcome-verified metric. Resolution rate is often measured more loosely, sometimes from the agent's own claim or a lack of repeat contact. Containment sits below both, measuring only whether a human stayed out of the call.

The bottom line

Containment tells you how many calls the agent kept, while task success tells you how many callers it actually helped. Track both and watch the gap between them, because that gap is where the contained-but-failed calls hide.

Want the honest outcome number on your own calls, not the flattering proxy? Book a demo and Evalgent will measure task success and containment side by side as your independent evaluator.

Related Articles