Test your voice agent
Why You Can't Trust Vendor-Reported Voice AI Metrics

Every voice AI vendor's site has a metrics section. Ninety-something percent accuracy, sub-second latency, high containment, a benchmark chart with the vendor's bar on top. It all looks like evidence, and buyers treat it that way — right up until the agent hits real callers and the numbers evaporate. The problem is not that vendors lie. It is that these numbers were never designed to help you compare vendors or predict production. This post explains why, and what to measure instead.
We will walk through the five reasons vendor metrics can't decide a purchase, the specific metrics that mislead most often, and how to replace them with numbers you can trust. It pairs with our vendor scorecard, which puts the alternative into practice.
Reason 1: the vendor measured its own metric
The first problem is simple: the vendor produced the number. It chose the test data, ran the evaluation, applied its own scoring, and published the result it liked. None of that is fraud — it is marketing, and marketing selects the flattering case. A number measured by the party trying to sell you something is not independent evidence, however precise it looks. This is the core reason independent evaluation exists, and why our voice agent evaluation guide insists the trustworthy number is the one you produce yourself.
Reason 2: the definitions don't match
Even when two vendors report the "same" metric, they rarely mean the same thing. "Accuracy" can mean word error rate on transcription, intent-classification accuracy, or task completion — three different numbers. "Latency" can mean model inference time, time to first audio, or full turn duration. "Containment" can count any call that avoided a human, including the ones where the caller gave up in frustration.
Because the definitions differ, the numbers do not compare. One vendor's 95% and another's 97% were measured against different targets, so the two-point gap is meaningless. Worse, the metric that sounds best is often the one that hides the failure — a high containment rate can describe an agent that traps callers, the trap our containment vs deflection guide describes.
Reason 3: the conditions were ideal
Vendor metrics are measured on clean audio, cooperative scripts, and the happy path — the conditions production is least likely to deliver. Real calls bring accents, background noise, interruptions, and callers going off-script, none of which a benchmark on studio-quality audio exercises. A vendor's clean-audio word error rate is a ceiling, not the accuracy you will see on a call from a moving car.
This is cherry-picking by construction: the reported number reflects the best-case slice, not the distribution your callers create. It is also why a great demo predicts nothing — a demo is the same curated success in live form, and the gap between it and production is exactly why voice agents fail once real traffic arrives.
Reason 4: the benchmark can't be reproduced
When a vendor cites a benchmark, you usually cannot rerun it. The exact dataset, preprocessing, and scoring script are rarely published in full, so the result is not reproducible — the standard that separates a measurement from a claim. Independent, production-scale voice benchmarks barely exist in this market, so most cited figures trace back to the vendor or to a source the vendor selected.
There is also a deeper trap. Once a metric becomes the number vendors compete on, they optimize for the benchmark rather than the outcome it stood for — Goodhart's law: when a measure becomes a target, it stops being a good measure. A vendor can top a public leaderboard and still fail your calls, because the leaderboard is what got optimized.
Reason 5: even the evaluator changes the answer
Here is the part buyers rarely account for: the scoring itself is not neutral. Run the same set of calls through two different evaluation setups — a different LLM judge, a different rubric, transcript-only versus audio-aware — and they can disagree sharply on which agent did better. The limits of transcript-based judging are covered in our LLM-as-judge piece.
That means a vendor's metric bakes in the vendor's choice of evaluator, which was also selected to flatter. Two honest teams can measure the same agent and reach different verdicts, so a number with no disclosed methodology is not comparable to anything — including the next vendor's number.
Vendor-reported vs your own metrics
The contrast is the whole argument.
| Aspect | Vendor-reported | Measured on your data |
|---|---|---|
| Who measured it | The vendor | You |
| Test data | Vendor's clean audio | Your real calls |
| Definition | Private, undisclosed | Fixed, identical across vendors |
| Conditions | Ideal, best-case | Accents, noise, edge cases |
| Reproducible | Rarely | Yes — you own the setup |
| Comparable across vendors | No | Yes |
The metrics that mislead most
A few numbers deserve special caution. A headline accuracy figure is usually clean-audio transcription, not task success, so it says nothing about whether calls got resolved. Average latency hides the tail, and the tail — p90 and p95 — is what callers actually feel; the comfortable one-way limit in the ITU-T G.114 telephony standard is 150ms, but an average can look fine while the slow calls feel broken. Containment counts absence of a human, not resolution. And any "N% better" claim without a disclosed, reproducible method is a marketing line, not a measurement. For regulated buyers, align what you measure to a recognized framework like the NIST AI Risk Management Framework rather than to vendor claims.
How to read a vendor metric without being misled
You cannot avoid vendor metrics, but you can interrogate them. Run any figure through these questions before you let it influence a decision.
1. Who measured it? — If the vendor produced the number, treat it as a claim, not evidence.
2. On what data? — Ask whether it was clean audio or realistic calls with accents and noise; a best-case dataset predicts a best case.
3. What is the exact definition? — Pin down what "accuracy," "latency," or "containment" means here, since each has several meanings that don't compare.
4. Is it the average or the tail? — For latency, insist on p90 and p95; an average hides the calls that feel broken.
5. Can you reproduce it? — If the dataset and scoring aren't disclosed, the number can't be verified or compared.
6. Does it measure the outcome? — Prefer task success over proxies like transcription accuracy or containment that say nothing about whether calls were resolved.
If a metric can't survive these questions, don't let it move your decision.
What to measure instead
The fix is not cynicism; it is measurement you own. Build one fixed set of scenarios and caller profiles that reflect your real traffic, run every vendor through the identical calls, and score each metric yourself with a single, disclosed definition. Weight the results by your priorities, test under load, and re-check over time. That is the vendor scorecard approach, and it is the only way to turn a wall of vendor numbers into a comparison you can defend. The pre-launch testing checklist covers the same discipline applied to a single agent.
Measuring voice AI honestly with Evalgent
Evalgent replaces vendor-reported metrics with numbers you own. Scenarios reproduce your real calls — accents, noise, interruptions, and adversarial callers — and run identically against every vendor or version you compare. Profiles vary caller conditions so no vendor can win on clean audio alone. Metrics are measured on your own data, with one fixed definition applied to everyone, so the results are finally comparable. Evaluations run the suite at concurrency, and Reviews let your team replay any call to hear what a score really represents, rather than trusting a figure with no methodology behind it.
The result is an evaluation you can defend: not the vendor's best-case number, but your own measurement on the calls you actually get. To measure the vendors you are weighing on identical calls, book a demo.
The bottom line
Vendor-reported voice AI metrics are self-measured, on cherry-picked data, with private definitions and unreproducible benchmarks — and even the choice of evaluator changes the answer. Numbers from different vendors do not compare, and none of them predict how an agent handles your callers.
Trust the number you measured yourself, on your own calls, scoring every vendor identically. A vendor's metric tells you what it wanted to show you; your measurement tells you what you will actually get.
Frequently asked questions
Why can't you trust vendor-reported voice AI metrics?
Because the vendor measured them — choosing the test data, the conditions, the definitions, and the scoring, then publishing the flattering result. Numbers from different vendors use different definitions and different data, so they don't compare, and the benchmarks behind them are rarely reproducible. Even the choice of evaluator changes the outcome, so only a metric you measured yourself is trustworthy.
Do voice AI vendors lie about their metrics?
Usually not. The problem is not dishonesty but selection: marketing reports the best-case number, measured on clean audio under ideal conditions with a favorable definition. That is legitimate marketing, not fraud. But a best-case figure produced by the seller cannot tell you how the agent performs on your real callers, which is why you have to measure it yourself on your own calls.
Why don't vendor accuracy numbers match production?
Because vendor accuracy is measured on clean, cooperative audio, and production brings accents, background noise, interruptions, and off-script callers. A clean-audio word error rate is a ceiling, not what you'll see on a real call. Headline accuracy also usually measures transcription, not task success, so a high number can coexist with calls that never actually got resolved.
What voice AI metrics are most misleading?
Headline accuracy, which usually measures clean-audio transcription rather than resolved tasks. Average latency, which hides the tail that callers actually feel. Containment, which counts any call that avoided a human, including frustrated hang-ups. And any "N% better" claim with no disclosed, reproducible method. Each can look strong while the agent fails the calls that matter to you.
How should you compare voice agent vendors fairly?
Build one fixed set of scenarios and caller profiles from your real traffic, run every vendor through the identical calls, and score each metric yourself with a single disclosed definition. Weight by your priorities and test under load. Running the same test set against every vendor, scored the same way, is the only route to a genuinely comparable, apples-to-apples result.
Are public voice AI benchmarks reliable?
Treat them cautiously. Most are not fully reproducible — the dataset, preprocessing, and scoring are rarely published in full — and once a benchmark becomes the number vendors compete on, they optimize for it rather than the outcome. A vendor can top a leaderboard and still fail your calls, so use public benchmarks as context, never as a substitute for measuring on your own data.
What does it mean that the evaluator changes the result?
It means scoring is not neutral. A different LLM judge, rubric, or transcript-versus-audio choice can rank the same agents differently on the same calls. So a vendor's metric embeds the vendor's chosen evaluator, itself selected to flatter. Two honest teams can measure one agent and disagree, which is why a number without a disclosed methodology can't be compared across vendors.
What should you measure instead of vendor metrics?
Measure task success — whether the caller's goal was accomplished — on your own calls, alongside accuracy weighted for critical entities, latency at p90 and p95, escalation accuracy, and a safety pass rate. Use one fixed definition for every vendor, test under load, and re-check over time. Own the measurement end to end, so the numbers reflect your callers rather than the vendor's best case.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more