Test your voice agent
How to Measure Task Success Rate for Voice Agents

# How to Measure Task Success Rate for Voice Agents
Quick answer
Task success rate is the share of calls where the caller's actual goal got done. You measure it by labeling each sampled call from the transcript, the audio, and your system of record, not the agent's own claim. It is the north-star outcome metric, separate from containment.
Most voice agent dashboards report the wrong number. They show how many calls the agent handled alone. They rarely show how many callers left with their problem solved. Those are different questions. A caller can stay on the line, hear a confident answer, and hang up with nothing fixed. That call looks like a win on the dashboard. It is a loss to the customer.
Task success rate closes that gap. It measures outcomes, not activity. This guide explains what the metric means, how to define success per call type, and how to label it from evidence. It also shows why an independent party should do the labeling. Evalgent measures task success on your real calls as the third-party evaluator.
Task success rate: the percentage of calls in which the caller's goal was actually completed, verified against evidence rather than the agent's own report.
What task success rate actually measures
Task success rate answers one question. Did the caller get what they called for? Nothing else counts toward the number. A pleasant tone does not count. A short call does not count. A confident closing line does not count.
The concept borrows from task analysis, the practice of breaking work into concrete goals. Each call type has a goal. The refund is issued or it is not. The appointment is booked or it is not. The address is updated or it is not. You score the goal, not the conversation around it.
Statisticians call this a success rate. It is a simple fraction. Divide completed goals by total attempts. The trick is not the math. The trick is defining the goal well and verifying it honestly.
Academic work on dialogue systems has treated task success as the anchor for decades. The PARADISE framework tied dialogue quality to whether the user's task was completed, as Walker and colleagues described in 1997. The idea predates modern voice agents. It still holds. Completion is the outcome that matters.
Our voice agent evaluation guide covers the wider metric set. Task success rate sits at the center of it. Everything else explains why success went up or down.
Task success rate is not containment
Teams confuse task success with containment constantly. The two look similar and mean opposite things. Getting this wrong flatters your agent and hides real failures.
Containment: the share of calls the agent handled without transferring to a human. Task success: the share of calls where the caller's goal got done.
A call can be contained and still fail. The agent never transferred, so containment counts it as a win. But the caller never got their refund. Task success counts it as a loss. High containment with low task success is a warning sign. It means callers are stuck with the agent and not getting helped.
The reverse also happens. A call transfers to a human who solves the problem. Containment marks it against the agent. Task success can still credit the outcome if the goal got done. That depends on how you scope the metric. Our containment versus deflection guide unpacks these definitions in detail.
Treat containment as an efficiency metric. Treat task success as an outcome metric. You want both, but you optimize for success. An agent that contains everything and solves nothing is worse than one that transfers often and solves problems.
Defining success criteria per call type
Success is not one definition. It changes with the reason for the call. A booking call and a billing dispute succeed in different ways. You need a written success criterion for each call type before you measure anything.
A good criterion is binary and checkable. It should read like a test, not a feeling. "Caller was satisfied" is not checkable. "Appointment created in the scheduler for the requested date" is checkable. Write criteria that a second reviewer could apply and reach the same verdict.
Some calls have partial outcomes. A caller asks two things and gets one resolved. Decide in advance how you score these. You can require full completion for a success. You can also score primary and secondary goals separately. Pick one rule and apply it consistently. Consistency matters more than the exact rule.
Edge cases need rules too. What about a caller who abandons mid-call? What about a request the agent is not allowed to fulfill? Mark these as defined categories, not silent gaps. Our scoring a voice agent conversation guide walks through building a rubric that reviewers can apply the same way every time.
Call type, success definition, and how to verify
The table below shows how success and verification change across common call types. The verification column is the important one. It names the evidence that proves the goal got done. These figures are illustrative examples, not benchmarks.
| Call type | Success definition | How to verify |
|---|---|---|
| Appointment booking | Slot reserved for the requested date and time | Booking exists in the scheduler system of record |
| Refund request | Refund issued to the correct account | Transaction posted in the billing platform |
| Address change | Account address updated to the new value | Field updated in the customer record |
| Order status check | Caller told the correct current status | Status in the transcript matches the order system |
| Password reset | Caller regained access to the account | Reset event logged in the identity system |
| Lead qualification | Required fields captured and routed | Complete record created in the CRM |
| Billing dispute | Dispute logged and next step confirmed | Case opened with correct details in the ticket system |
Notice a pattern. Verification almost always points to a system outside the call. The transcript alone is not enough. The agent saying "your refund is processed" is a claim. The refund appearing in the billing platform is proof. That distinction drives the next section.
Where the label comes from
The single most common mistake is trusting the agent's own claim. Agents are built to sound helpful. They say "all set" and "you are good to go." Those phrases are not evidence. They are the thing you are supposed to be checking.
Labeling task success needs three sources together. Each catches errors the others miss.
- The transcript shows what was said and whether the caller's goal was even understood. It reveals misroutes, missed intents, and wrong information.
- The audio shows what the transcript hides. Tone, long silences, talk-over, and frustration live in the audio. A caller can say "yes, thanks" through gritted teeth.
- The system of record shows whether the action actually happened. This is the ground truth. It confirms or contradicts the agent's claim.
When these three disagree, you have found a false success. The agent claimed a booking. The transcript sounds fine. The scheduler has no appointment. That is a failed call that every naive metric would have counted as a win. Our transcript versus audio evaluation guide explains why you need both signals, not one.
This is also where independence matters. A vendor grading its own agent has an incentive to read ambiguous calls generously. An independent evaluator applies the same rubric without that pressure. The point of a third party is not distrust. It is that the same standard gets applied to every call.
How to measure task success rate step by step
Follow these steps to produce a task success rate you can defend. Evalgent runs this process on your real calls as the independent evaluator, but the method is the same whoever applies it.
1. List your call types. Group calls by the caller's goal. Start with the highest-volume types. You do not need every type on day one.
2. Write a success criterion for each type. Make each one binary and checkable. Name the exact evidence that proves completion.
3. Choose your evidence sources. Confirm you can access the transcript, the audio, and the relevant system of record for each call type.
4. Draw a sample. Pull a random set of calls per type. Random selection keeps the sample fair and representative.
5. Label each call against its criterion. Check all three sources. Mark success, failure, or a defined edge category. Never rely on the agent's claim alone.
6. Calculate the rate per type and overall. Divide successes by labeled calls. Report each call type separately, then a weighted total.
7. Attach a confidence interval. A sample gives an estimate, not an exact figure. State the range so no one over-reads a small sample.
8. Diagnose the failures. Read the failed calls. Group them by cause. This is where the metric turns into fixes.
The output is a number per call type and a reason behind every failure. You can benchmark this on your own data rather than trusting generic scores.
Sampling versus full measurement
You do not have to label every call. For most teams, that would be slow and expensive. A well-drawn sample gives a reliable estimate at a fraction of the cost. This is the logic of statistical sampling).
The key rules are simple. Draw the sample at random. Cover each call type. Make the sample large enough to give a tight interval. A larger sample narrows the range around your estimate. A tiny sample can swing wildly and mislead you.
Full measurement has its place. Automated labeling can screen every call for likely failures. You then human-review a sample to check the automated labels. This blends coverage with accuracy. The automated pass finds candidates. The human pass confirms the truth.
Think about the trade-off in terms of precision and recall. Precision is how often a flagged failure is really a failure. Recall is how many real failures you catch. An automated screen tuned for recall catches most problems. A human sample restores precision. You want both, so you combine them.
Whatever you choose, report the method with the number. A task success rate without its sample size and method invites misuse. Say how many calls, drawn how, labeled by whom. Transparency about method is part of the metric.
Why task success rate is the north-star metric
A north-star metric is the one number that best tracks the value you deliver. For a voice agent, that is task success rate. It is the key performance indicator that maps directly to customer outcomes. Other metrics explain it. This one defines it.
Latency, containment, and sentiment all matter. But they are inputs, not outcomes. A fast agent that fails tasks is a fast failure. A contained agent that solves nothing wastes the caller's time. Task success rate keeps the whole team pointed at the outcome that customers feel.
It also resists gaming. You can inflate containment by refusing transfers. You can lower latency by cutting corners. You cannot fake a refund appearing in the billing system. Because task success is verified against the system of record, it is hard to game and easy to trust. That is exactly what a north-star metric should be.
Build your dashboard around it. Put task success rate at the top, split by call type. Rank your other metrics as explanations underneath. Our voice agent metrics scorecard shows how to lay this out so the outcome leads and the diagnostics follow.
Frequently asked questions
What is task success rate for a voice agent?
Task success rate is the percentage of calls where the caller's goal was actually completed. It is measured against evidence, not the agent's own claim. A refund counts as a success only when it appears in the billing system. It is the clearest outcome metric for a voice agent.
How is task success rate different from containment?
Containment measures the share of calls handled without a human transfer. Task success measures whether the caller's goal got done. A call can be contained and still fail, because the agent kept the caller but solved nothing. Optimize for task success and treat containment as an efficiency measure alongside it.
How do you define success for a voice agent call?
Write a binary, checkable criterion for each call type. It should read like a test a second reviewer could apply. For a booking, success means an appointment exists in the scheduler for the requested date. Name the exact evidence that proves completion, and set rules for partial outcomes in advance.
Why can't you trust the agent's own claim of success?
Voice agents are built to sound helpful and confident. Saying "your refund is processed" is a claim, not proof. The refund appearing in the billing system is proof. When the agent's claim and the system of record disagree, you have found a false success that naive metrics would count as a win.
How do you label task success from a transcript?
The transcript alone is not enough. Read it with the audio and the system of record together. The transcript shows what was said. The audio shows tone and frustration. The system of record shows whether the action actually happened. A call succeeds only when all three sources confirm the goal was met.
Should you sample calls or measure every one?
Sampling is enough for a reliable estimate for most teams. Draw calls at random, cover each call type, and use a large enough sample to keep the confidence interval tight. Automated screening can scan every call for likely failures, with a human-reviewed sample to confirm the labels and restore precision.
Who should measure voice agent task success?
An independent evaluator should apply the rubric. A vendor grading its own agent has an incentive to read ambiguous calls generously. A third party applies the same standard to every call without that pressure. Evalgent does this on your real calls, so the number reflects outcomes rather than the vendor's framing.
What sample size do you need for task success rate?
There is no single number, but bigger samples give tighter estimates. A small sample can swing widely and mislead you. Always report the sample size, how it was drawn, and who labeled it, alongside a confidence interval. The method is part of the metric, and hiding it invites misuse of the result.
The bottom line
Task success rate is the only voice agent metric that measures whether the caller's goal actually got done. Measure it from the transcript, the audio, and your system of record, and have an independent party apply the same standard to every call.
Measure task success on your real calls
Evalgent is the independent, third-party evaluator that scores task success on your own recorded calls. Ready to see your real task success rate, verified against your systems of record? Book a demo and we will measure it on your own calls.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more