Evalgent
Back to Blog
Voice AI Evaluation

A Contact Center Manager's Guide to Voice Vendors

Deepesh Jayal
12 min read
A Contact Center Manager's Guide to Voice Vendors

# A contact center manager's guide to voice vendors

Quick answer

A contact center manager choosing a voice agent vendor should judge each vendor on real containment, average handle time, behavior at peak concurrency, and clean handoff to human agents. Test all four on your own calls, at your own volume, with an independent evaluator. Demos do not predict operational fit.

What the operations lens looks for

You run the floor. Every vendor pitch lands on your desk as a promise about numbers you are already measured on. Containment. Average handle time. Service level. Occupancy. The vendor's slides show all of them improving. Your job is to find out whether that holds when 300 callers arrive at once and the agent must hand three of them to a live rep mid-sentence.

This guide is written for the operations lens. Not the customer experience view, and not the technology stack view. It covers what to measure, how to run the evaluation, and the red flags that separate a demo win from a floor that runs cleaner. It sits alongside the pillar on how to evaluate voice agent vendors and the deeper mechanics in our voice agent evaluation guide.

Why the operations lens is different

A CX leader asks whether callers feel heard. A CTO asks whether the stack scales and stays secure. You ask a narrower, harder question. Does this change the shape of my day?

The contact center is a queueing system. Calls arrive, wait, get served, or abandon. Every metric you own flows from that. A voice agent is a new server in the queue. It behaves differently from a human one. It never gets tired. It also fails in ways a human never would, all at once, under load.

So the operations question is not "is it good." The question is "what does it do to my queue, my staffing plan, and my handoffs." That is measurable. It is also the part vendors test least.

Containment, not deflection

The first number to pin down is containment. Vendors love to blur it with deflection. The two are not the same.

> Containment: the share of calls the voice agent fully resolves without any human involvement, where the caller's goal was actually met. Not merely calls that avoided a transfer.

> Deflection: the share of calls steered away from a human queue, whether or not the caller's problem was solved.

A deflected call that did not resolve comes back. It returns as a repeat call, a worse first call resolution rate, and an angrier caller. So a high deflection number can hide a containment problem. We unpack the trap in the guide on containment versus deflection.

Insist on true containment measured on resolved outcomes. Define "resolved" before you test. Then check whether contained calls stay contained, by tracking repeat contacts over the next few days. If callers return for the same reason, the agent deflected. It did not contain.

Average handle time, and the trap inside it

Handle time is where a voice agent can help you most, and mislead you most.

> Average handle time (AHT): the mean duration of a caller interaction, including talk time and any wrap-up work, across a defined period.

A voice agent has no after-call work in the human sense. That alone can pull reported AHT down. But the metric that matters to you is blended. What is the AHT across the whole queue once the agent handles some calls and hands the rest to people?

A fast agent that escalates poorly can raise your true AHT. The human now inherits a caller who is already frustrated. That caller has to be re-briefed from scratch, so the call runs long. Measure AHT in three parts: agent-only calls, human-only calls, and the escalated calls that touch both. The last group is where hidden cost lives.

Tie this back to service level and occupancy. A voice agent shifts your arrival pattern to the human queue. It removes the simple, short calls and leaves the hard, long ones. Your human occupancy math changes. Plan for it.

Behavior at peak concurrency

Here is the failure mode that demos never show. A voice agent looks flawless at one call. It can degrade badly at 200.

Under load, latency climbs. The pause before the agent speaks grows. Turns collide. The agent starts talking over callers, or goes silent, or drops context between turns. Some of these faults appear only above a concurrency threshold. Below it, everything looks fine.

This matters because contact center traffic is spiky. Traffic engineering has modeled this for a century, from Erlang C staffing math to modern workforce management. Your busy hour is not your average hour. A product recall, a weather event, or a billing run can triple arrivals in minutes. The vendor's clean demo told you nothing about that moment.

So test at your real peak, not your average. Push past it. Watch what breaks first, and how it recovers. Our guide on latency in voice agents covers what to watch as concurrency rises. This is exactly the kind of work covered in stress testing voice AI.

Clean handoff to human agents

No voice agent handles everything. The handoff is where operations wins or loses.

There are two kinds. A warm handoff passes the caller to a human with full context. A cold handoff dumps the caller into a queue with nothing. Cold handoffs generate complaints and repeat the whole conversation.

> Warm handoff: the voice agent transfers the caller to a human agent together with the transcript, the intent, and the data already collected, so the human resumes without re-asking.

Test three things on every handoff. Does the caller reach the right skill or queue? Does the human receive the context, or start blind? Does the escalation trigger at the right moment, not after the caller has repeated themselves five times? The escalation guide walks through the trigger logic in detail.

A clean handoff protects your first call resolution and your handle time at the same time. A broken one quietly taxes both.

After-hours and overflow coverage

Two operational jobs suit voice agents well. After-hours coverage, and overflow during spikes. Both are worth buying for. Both need testing under their real conditions.

After-hours means the agent works with no human safety net behind it. If it cannot resolve or safely take a message, the caller is stranded until morning. So the after-hours escalation path needs a fallback that does not depend on a live queue. Test what happens when there is no one to transfer to.

Overflow means the agent absorbs the top of a spike so humans are not buried. That is a concurrency test and a handoff test at once. Confirm the agent scales up fast and hands back gracefully as the spike passes.

Reporting and QA fit

A voice agent that cannot report into your existing tools creates blind spots. You lose the single view of the floor.

Ask how the vendor exposes call data. Can you pull agent transcripts, containment, AHT, and escalation reasons into your reporting? Can your QA team score voice agent calls on the same scorecard they use for humans? If the agent lives in a separate dashboard with its own definitions, your numbers stop reconciling.

Insist on shared definitions and exportable data. Your QA and workforce management processes should treat the voice agent as one more agent group, measured the same way.

Operational criteria, what to measure, and red flags

Use this as your evaluation scorecard. Score every vendor on the same rows, using your own call data.

Operational criterionWhat to measureRed flag
True containmentResolved-without-human rate, plus repeat contacts over 3–7 daysVendor quotes deflection and calls it containment
Average handle timeAgent-only, human-only, and escalated-call AHT, measured separatelyOnly agent-only AHT shown, blended number hidden
Peak concurrencyLatency, turn collisions, and dropped context at and above your busy hourDemo runs one call at a time, no load numbers
Warm handoffRight-queue routing, context passed, correct escalation timingCold transfers, caller re-asked everything
After-hours coverageResolution and safe fallback with no human queue behind the agentFallback assumes a live agent is always available
Reporting and QA fitExportable data on the same definitions your team already usesSeparate dashboard, vendor-only metric definitions

How to run an operations evaluation of a voice agent vendor

Run this process the same way for every vendor, so the comparison is fair. It is the operations version of a structured bake-off.

1. Define resolved and set your metric targets. Write down what counts as a contained, resolved call for each call type. Set targets for containment, blended AHT, service level, and escalation rate before any vendor sees the test.

2. Build a test set from your real calls. Pull a representative mix of your actual call reasons, accents, noise, and edge cases. Include the hard calls, not just the easy ones. This becomes your shared benchmark, as detailed in benchmark voice agents on your own data.

3. Run identical calls on every vendor. Put each vendor through the same scenarios so differences come from the agent, not the test. The discipline is covered in compare voice agents on the same test cases.

4. Test at and above your peak concurrency. Recreate your busy hour, then push higher. Record latency, turn behavior, and context retention as load climbs.

5. Score every handoff. For each escalated call, check routing, context transfer, and trigger timing. Count how many callers had to repeat themselves.

6. Measure blended AHT and service level. Model the effect on the human queue once simple calls are removed and hard ones remain. Update your staffing plan against that new mix.

7. Confirm reporting and QA integration. Verify you can export the data and score calls on your own scorecard, not the vendor's.

8. Decide on measured results, and keep auditing. Choose on the numbers you produced. Then re-run the evaluation periodically, because a model update can quietly change behavior.

Change management for your staff

The technology decision is half the work. The floor is the other half.

Adding a voice agent changes the human job. The easy calls leave. What remains is harder, longer, and more emotional. Your handle time targets for humans should rise, not fall. Your coaching should shift toward the complex work the agent hands over.

Tell your team early what the agent does and does not handle. Show them how escalations arrive, and what context they will get. Update QA scorecards so humans are judged on the new call mix, not the old one. A voice agent rollout that surprises the floor erodes trust and gets worked around.

Where independent evaluation fits

Every check in this guide has the same weakness. If the vendor runs the test, the vendor controls the conditions. Containment gets redefined. Load tests use the vendor's easy scenarios. Handoff quality is scored on the vendor's terms.

Evalgent is an independent, third-party evaluator for AI voice agents. We run your calls, at your volume, and measure what your floor actually cares about. We test containment on resolved outcomes, blended handle time, behavior at peak concurrency, and whether every handoff arrives warm. Because we do not sell the agent, the numbers are not marketing. They are evidence you can put in front of your leadership and your team. That premise is explained in independent voice AI evaluation.

If you also own the experience or the stack, the sibling guides for the CX leader, the CTO, the VP of engineering, the procurement lead, and the compliance officer cover the same decision from their angles.

Frequently asked questions

What should a contact center manager look for in a voice agent vendor?

A contact center manager should look for true containment on resolved outcomes, blended average handle time, stable behavior at peak concurrency, and clean warm handoffs to human agents. Reporting must fit existing QA and workforce tools. The strongest signal is measured performance on your own calls, not a demo or a vendor-reported figure.

What is the difference between containment and deflection for voice agents?

Containment is the share of calls the agent fully resolves without a human, where the caller's goal was met. Deflection is the share steered away from a human queue, whether or not the problem was solved. Deflected but unresolved calls return as repeat contacts, so a high deflection number can hide a real containment problem.

How do you test a voice agent at peak call concurrency?

Recreate your busy-hour arrival volume, then push above it. Run many simultaneous calls and watch latency, turn collisions, silences, and dropped context. Some faults appear only above a concurrency threshold. Vendor demos usually run one call at a time, so peak-load testing is the only way to see how the agent degrades under real spikes.

What is a good average handle time for an AI voice agent?

There is no universal target, because handle time depends on call type and complexity. What matters is blended average handle time across the whole queue, measured in three parts: agent-only calls, human-only calls, and escalated calls that touch both. A fast agent that escalates poorly can still raise your true handle time by re-briefing frustrated callers.

How should a voice agent hand off calls to human agents?

A voice agent should perform a warm handoff. It routes the caller to the correct skill or queue, passes the transcript, intent, and collected data, and triggers at the right moment. The human then resumes without re-asking anything. Cold handoffs, which drop the caller into a queue with no context, generate complaints and repeat the whole conversation.

Can voice agents cover after-hours and overflow call volume?

Voice agents suit after-hours coverage and spike overflow well. Both need testing under real conditions. After-hours work has no human safety net, so the fallback cannot assume a live queue exists. Overflow requires fast scale-up and graceful handback as the spike passes. Test the escalation path when there is genuinely no one available to receive a transfer.

How does independent evaluation help contact center vendor selection?

Independent evaluation removes the vendor's control over test conditions. A third party like Evalgent runs your real calls at your real volume and measures containment, blended handle time, concurrency behavior, and handoff quality on consistent definitions. Because the evaluator does not sell the agent, the results are evidence rather than marketing, and they compare vendors fairly on the same test set.

How do you manage change when adding a voice agent to your team?

Manage change by preparing the floor early. The agent removes easy calls, leaving harder, longer ones, so raise human handle time targets and shift coaching toward complex work. Show agents how escalations arrive and what context they receive. Update QA scorecards for the new call mix. A rollout that surprises staff erodes trust and gets worked around.

The bottom line

Choose a voice agent vendor on measured containment, blended handle time, peak-concurrency behavior, and clean handoff, tested on your own calls. Let an independent evaluator run those tests, because a vendor that controls the conditions controls the result.

Ready to see how a vendor holds up at your real call volume? Book a demo and we will run your calls, measure what your floor cares about, and hand you the evidence.

Related Articles