Test your voice agent
How to Evaluate Voice Agent Vendors: A 2026 Scorecard

Enterprises rarely bet on a single voice agent vendor anymore. Many run four or five at once — different vendors for different use cases, or several in parallel to hedge against vendor lock-in and outages. That makes vendor evaluation a recurring, high-stakes decision, and it is one most teams make on the worst possible evidence: a polished demo and a page of vendor-reported metrics. This guide replaces that with a scorecard you can defend — the same criteria, scored on your own calls, for every vendor you consider.
We will cover why vendor-reported numbers can't decide this for you, the seven criteria that belong on the scorecard, and a step-by-step process for running the evaluation. It builds on the discipline in our voice agent evaluation guide and the pre-launch testing checklist, applied to the vendor-selection decision.
Why you can't evaluate vendors on their own numbers
Almost every metric a vendor publishes is measured by the vendor, on data the vendor chose, under conditions the vendor controlled. That is not dishonesty — it is marketing — but it makes cross-vendor comparison meaningless. One vendor's "97% accuracy" and another's "95%" were measured on different audio, different tasks, and different definitions of accuracy, so the numbers do not compare.
The deeper problem is that independent, production-scale benchmarks barely exist in this market. Even the choice of evaluator matters: run the same set of calls through two different evaluation setups and they can disagree sharply on which agent performed better. The only number you can trust is the one you produced yourself, on your own scenarios, scoring every vendor identically. That is the entire premise of the scorecard — and the reason a demo, which is a curated success under ideal conditions, tells you nothing about how a vendor handles the callers you actually get.
The voice agent vendor scorecard
Score every vendor on the same seven criteria — an adaptation of established software quality thinking to the realities of a phone call. The weights are yours to set — a healthcare intake agent weights safety and compliance far higher than an outbound sales agent — but the criteria stay constant so the comparison is fair.
| Criterion | What it measures | Typical weight |
|---|---|---|
| Accuracy | Transcription and understanding on your audio, weighted for critical entities | High |
| Latency | Time to first audio and turn latency at p90/p95 | High |
| Task success | Whether the caller's goal was actually accomplished | Highest |
| Escalation & recovery | Correct handoffs and graceful recovery from errors | High |
| Safety & compliance | Refusals, data protection, and policy adherence | Highest for regulated |
| Scalability | Accuracy and latency held under peak concurrency | Medium–High |
| Cost | Total cost per resolved outcome, not just per minute | Medium |
Accuracy is transcription and intent on your callers — accents, noise, and the names, dates, and amounts your calls contain — not a vendor's clean-audio word error rate. Latency is the pause the caller feels, measured at the tail; human sensitivity to delay is well established, with the ITU-T G.114 standard putting the comfortable one-way limit at 150ms. Task success is the criterion that outranks the rest, because an agent can transcribe perfectly and still fail the call. Escalation and recovery decide whether a stuck caller reaches a human or gets trapped. Safety and compliance should map to a recognized framework such as the NIST AI Risk Management Framework, and matter most in regulated industries. Scalability re-checks the top criteria under load. Cost is measured per resolved outcome, the angle our voice agent cost guide unpacks — the cheapest per minute is often not the cheapest per solved problem.
Weighting the scorecard for your use case
The criteria stay fixed, but the weights are where the scorecard becomes yours. A single set of weights applied to every deployment hides the tradeoffs that actually decide the winner.
For a customer support agent, task success and containment lead, with latency close behind, because a slow or unresolved call drives callers back to a human. For a healthcare intake or financial services agent, safety and compliance dominate — a single mishandled disclosure outweighs a small latency edge. For an outbound sales agent, latency and conversational naturalness rise, since a stilted or laggy opener loses the call in seconds.
Write the weights down before you score, and tie them to the outcomes your business cares about. Keep them in the record so procurement can see why one vendor won. When the weights are explicit, the decision stops being a matter of taste and becomes a matter of evidence — which is what a defensible RFP requires.
How to evaluate voice agent vendors
The process is what makes the scorecard fair. Run it identically for every vendor.
1. Define the use case and success criteria — Write down the calls the agent must handle and what "resolved" means for each, before you look at any vendor.
2. Build one shared test set — Assemble a fixed set of scenarios and caller profiles — accents, noise, interruptions, edge cases — that every vendor will face.
3. Run the same calls on every vendor — Put each vendor through the identical scenarios, so differences come from the agent, not the test, as our same test cases discipline requires.
4. Score on your own data — Measure each criterion yourself from the recorded calls, rather than accepting vendor-reported figures.
5. Weight by your priorities — Apply the weights your use case demands, so safety-critical or latency-critical needs dominate the total where they should.
6. Test at production scale — Re-run the critical criteria at your expected peak concurrency, since both accuracy and latency degrade under load.
7. Decide on the total, and keep auditing — Choose on the weighted score, then re-run the scorecard periodically, because a vendor that wins today can regress after a model update.
Common mistakes when evaluating vendors
The errors are predictable. Trusting the demo, which is engineered to succeed. Comparing vendor-reported metrics that were measured differently. Testing on the vendor's sample data instead of your own calls. Weighting cost too heavily and discovering the cheap agent escalates or fails more, raising the true cost per resolved call. Evaluating once and never re-auditing, so a silent regression after a model update goes unnoticed. And scoring only the happy path, so the first angry or off-script caller finds the failure in production — the pattern behind why voice agents fail in production. Each is avoidable with a shared test set scored on your own data.
Demo evaluation vs scorecard evaluation
The contrast is why the scorecard exists.
| Dimension | Vendor demo | Scorecard evaluation |
|---|---|---|
| Data | Vendor's clean audio | Your real calls |
| Conditions | Ideal, curated | Accents, noise, edge cases |
| Metrics | Vendor-reported | Measured by you, identically |
| Coverage | Happy path | Failures and load |
| Comparability | None across vendors | Direct, same test set |
| Output | "It looked great" | A weighted, defensible score |
Evaluating voice agent vendors with Evalgent
Evalgent is built to run exactly this scorecard, independently of any vendor. Scenarios reproduce your real calls — happy paths, edge cases, interruptions, and adversarial callers — and run identically against every vendor you are comparing. Profiles vary caller accent, pace, and line quality across those calls, so a vendor cannot win on clean audio alone. Metrics measure each scorecard criterion — accuracy, latency percentiles, task success, escalation, safety — on your own data, against the weights you set, so the result is one comparable score per vendor. Evaluations run the whole suite at concurrency, and Reviews let your team replay any call to hear why one vendor scored higher than another. Because the same test set runs against everyone, the comparison is finally apples to apples.
The result is a vendor decision you can defend to procurement and security. Not "the demo looked good," but "we scored every vendor on the same calls, and here is the evidence." To set up an independent evaluation of the vendors you are weighing, book a demo. For the discipline behind the scores, see testing vs evaluation and the LLM-as-judge limits that shape how the scoring is done.
The bottom line
Evaluate voice agent vendors on the same criteria, scored on your own calls, not on the numbers each vendor reports about itself. Accuracy, latency, task success, escalation, safety, scalability, and cost — weighted for your use case — give you one comparable score per vendor.
Run identical scenarios against every candidate, measure the results yourself, and re-audit over time, because the winner can change after a model update. A demo tells you a vendor can succeed; the scorecard tells you which one will, on the callers you actually get.
Frequently asked questions
How do you evaluate a voice agent vendor?
Evaluate a voice agent vendor by scoring it on fixed criteria — accuracy, latency, task success, escalation, safety and compliance, scalability, and cost — using your own test calls rather than vendor-reported numbers. Run the same scenarios against every vendor you are comparing, measure each criterion yourself from the recorded calls, weight by your priorities, and decide on the total weighted score.
Why can't you trust vendor-reported voice AI metrics?
Because each vendor measures its own metrics, on data it selected, under conditions it controlled, and with its own definitions. One vendor's accuracy figure and another's were produced differently, so they do not compare. Independent production-scale benchmarks are scarce, and even the choice of evaluator changes the result. The only trustworthy number is one you produced yourself, scoring every vendor identically.
What should be on a voice agent vendor scorecard?
A voice agent vendor scorecard should score accuracy, latency at the tail, task success, escalation and recovery, safety and compliance, scalability under load, and cost per resolved outcome. Keep the criteria constant across vendors so the comparison is fair, and set weights by use case — safety dominates in regulated industries, latency in real-time sales — so the total reflects what matters to you.
How do you compare voice agents on the same test cases?
Build one fixed set of scenarios and caller profiles — accents, noise, interruptions, edge cases — and run every vendor through the identical calls. Score each criterion yourself from the recordings, so differences come from the agent rather than from different tests. Running the same test set against every vendor is the only way to get a genuinely comparable, apples-to-apples result.
How often should you re-evaluate a voice agent vendor?
Re-evaluate on a regular cadence, not just at selection, because a vendor's performance changes when it updates its models. Re-run the scorecard periodically and after any known vendor change, and sample live calls continuously so regressions surface between full evaluations. Continuous auditing keeps a vendor honest over the life of the contract, rather than trusting a single passing result from procurement time.
Should cost be the deciding factor when choosing a vendor?
No. Cost matters, but measured per resolved outcome, not per minute. A cheaper-per-minute agent that escalates or fails more often can cost more per solved problem and frustrate more callers. Weight cost alongside task success, accuracy, and safety, and treat the lowest total cost per resolved call — not the lowest headline rate — as the real economic comparison between vendors.
How do you evaluate voice agents for a regulated industry?
Weight safety and compliance highest, and map the checks to a recognized framework so the evaluation is defensible to security and legal. Verify data protection, policy adherence, and correct refusals directly on your own calls, including adversarial ones. Confirm the vendor holds accuracy and safety under load, and document the scored results, since regulated procurement typically requires evidence rather than vendor assurances.
What's the difference between a demo and a real vendor evaluation?
A demo is a curated success on the vendor's clean audio and happy path, engineered to look good. A real evaluation runs your own calls — accents, noise, interruptions, failures — identically against every vendor and scores the results yourself. The demo shows a vendor can succeed under ideal conditions; the evaluation shows which vendor actually will, on the callers you cannot control.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more