Evalgent
Back to Blog
Voice AI Evaluation

How to Compare Voice Agents on the Same Test Cases

Deepesh Jayal
12 min read
How to Compare Voice Agents on the Same Test Cases

Ask two vendors to prove they're better and each will hand you a different test — the one their agent wins. That is why "our agent scored 95%" tells you nothing: the score and the test came from the same source. The only comparison that means anything runs the same calls against every agent, scored the same way. This guide is about building that shared test set from your real traffic, running it identically, and reusing it so a one-time comparison becomes a permanent asset.

We will cover what a shared test set actually contains, how to build one from your own calls, how to run it fairly against every agent, and how to reuse it over time. It is the practical companion to our vendor benchmarking guide and the vendor scorecard.

Why "the same test cases" is the whole game

A comparison is only as fair as what varies between the things compared. If Agent A and Agent B face different calls, different audio, or different scoring, any difference in their scores could come from the test rather than the agent. Hold the test cases constant and the only variable left is the agent — which is the entire point.

This is why a vendor's own numbers can't be compared across vendors: each was produced on a different test, so the numbers measure different things, the problem our vendor metrics piece unpacks. A shared test set removes that ambiguity. Run the identical scenarios, on the identical caller profiles, scored by the identical rules, and a five-point gap is a real five-point gap — not an artifact of two different exams.

What a shared test set contains

A test set is more than a list of questions. Three parts make it usable across agents.

Scenarios — the calls the agent must handle, each a defined situation: a caller changing an order, disputing a charge, asking something out of scope. Cover the happy path and, deliberately, the hard cases.

Caller profiles — the how of each call: accent, speaking pace, background noise, and behavior, from cooperative to confused to hostile. The same scenario across several profiles is where agents separate.

Expected outcomes — what "resolved" means for each scenario, written down in advance as assertions: the right action taken, the correct information given, an escalation when required. Without these, scoring drifts into opinion.

Together these make the set portable: any agent can be dropped into the same scenarios, met by the same callers, and judged against the same outcomes.

How to build a shared test set from your calls

The set is worth building once, carefully, because you will reuse it for years.

1. Mine your real calls — Start from actual traffic, not imagined dialogues, so the scenarios reflect what callers really do.

2. Curate representative scenarios — Select the common paths and the costly edge cases; a set that only covers the happy path compares agents on the easy 20%.

3. Add the hard and adversarial cases — Include noise, interruptions, off-script callers, and attempts to break the agent, since that is where agents diverge.

4. Write expected outcomes — Define "resolved" for each scenario as explicit assertions, so scoring is objective and repeatable.

5. Build a range of caller profiles — Vary accent, pace, and line quality, and run every scenario across them, using synthetic callers to reach the range at scale.

6. Version the set — Store it as a fixed, versioned artifact so every agent is measured against the exact same cases, and changes are tracked.

Sourcing real calls responsibly

Building the set from real calls raises an obvious question: how do you use production audio without exposing sensitive data? Treat the test set like any other sensitive asset. Redact or mask personal information as you curate calls, keep the set access-controlled, and prefer synthetic callers that reproduce the pattern of a hard call — the accent, the interruption, the disputed charge — without carrying a real customer's private details. For regulated workflows, align this handling with a recognized framework such as the NIST AI Risk Management Framework, so the test set itself passes a security review. Done well, you capture what makes your calls hard without turning the suite into a data-leak risk.

Running the set identically against every agent

Building the set is half the work; running it fairly is the other half. Replay the identical calls against each agent, holding everything constant except the agent under test — same scenarios, same profiles, same order, same scoring rules. Any variable you let drift becomes a confound that muddies the comparison.

Score each agent against the pre-written expected outcomes, and measure every layer — accuracy, task success, latency at the tail against a threshold like the ITU-T G.114 150ms comfort limit, escalation, and safety. Run the set at your expected concurrency, since rankings can invert under load. And keep the result reproducible: if rerunning the same set produces the same ranking, the comparison is signal; if it doesn't, it was noise.

Reusing the set over time

The payoff of a shared test set is that it does not expire. Once built, it becomes a regression suite: rerun it whenever you change a prompt, swap a model, or a vendor ships an update, and a drop from the previous run flags a regression before callers find it. It is also the fair basis for A/B testing two versions of your own agent, since both face the identical cases.

Feed new failure modes back into the set as production surfaces them, so it grows more representative over time. A vendor that led at selection can regress after a model update, and the same set that chose it will catch the slip — the ongoing discipline behind the pre-launch testing checklist.

Ad-hoc testing vs a shared test set

The contrast is why the shared set is worth the effort.

AspectAd-hoc testingShared test set
Calls per agentDifferent each timeIdentical for all
ScoringBy feelPre-written outcomes
Comparable across agentsNoYes
ReusableNoYes — a regression suite
Catches regressionsNoYes, on every change
Reflects your trafficRarelyBy construction

Common mistakes

The errors are familiar. Letting each vendor bring its own test, so nothing compares. Building the set from imagined dialogues instead of real calls. Covering only the happy path, so agents look identical until production. Scoring by impression rather than pre-written outcomes, so the verdict shifts with the reviewer. Changing the set between agents, which quietly confounds the comparison. And building the set once for a bake-off, then never reusing it — throwing away the regression suite you already paid to create. Each is avoided by treating the test set as a fixed, versioned, reused artifact.

Comparing voice agents with Evalgent

Evalgent is built around the shared test set. Scenarios capture your real calls — happy paths, edge cases, interruptions, and adversarial callers — as a fixed, versioned suite. Profiles vary caller accent, pace, and line quality, so every agent faces the same range, not a flattering slice. Metrics score each agent against pre-defined expected outcomes with one fixed definition, so the results are directly comparable. Evaluations replay the identical set against every agent or version at concurrency, and Reviews let your team hear any call behind a score. Because the same suite runs against everyone and can be rerun on demand, it doubles as your regression gate long after the initial comparison.

The result is a comparison you can trust and keep using: the same calls, scored the same way, for every agent — then reused on every change. To compare the agents you are weighing on one identical test set, book a demo.

The bottom line

Comparing voice agents fairly means one thing: the same test cases, run identically against every agent, scored against outcomes you defined in advance. Anything else compares different exams and calls the result a winner.

Build the set once from your real calls, version it, and reuse it — as a bake-off today and a regression suite forever after. The vendor with the best demo brought its own test; the agent that wins on your shared test set is the one that will hold up on your calls.

Frequently asked questions

How do you compare voice agents on the same test cases?

Build one fixed test set from your real calls — scenarios, caller profiles, and expected outcomes — and run it identically against every agent, holding calls, conditions, and scoring constant. Score each agent against the pre-written outcomes and measure every layer. Because only the agent varies, differences in the results are attributable to the agent rather than to different tests.

What goes into a voice agent test set?

Three parts: scenarios, the calls the agent must handle including edge and adversarial cases; caller profiles, varying accent, pace, noise, and behavior; and expected outcomes, an explicit definition of "resolved" for each scenario. Together they make the set portable, so any agent can be dropped into the same cases, met by the same callers, and judged against the same objective outcomes.

Why can't you compare vendors on their own test results?

Because each vendor produced its result on its own test — the one its agent wins — with its own scoring. Numbers from different tests measure different things and don't compare. Only running the same calls against every vendor, scored the same way, removes that ambiguity, which is why a shared test set built from your own traffic is the basis of a fair comparison.

How do you build a voice agent test set from real calls?

Mine actual traffic rather than imagining dialogues, curate the common paths and the costly edge cases, and deliberately add noise, interruptions, and off-script callers. Write explicit expected outcomes for each scenario, build a range of caller profiles, and run every scenario across them with synthetic callers. Store the whole thing as a fixed, versioned artifact so every agent faces identical cases.

How many test cases do you need to compare voice agents?

Enough to cover your real scenarios and enough calls per scenario that results are stable rather than a fluke of a few runs. Coverage matters as much as count: include the edge and adversarial cases where agents actually diverge, and run each scenario across several caller profiles. If rerunning the set produces the same ranking, you have enough; if it doesn't, add more.

Can you reuse a voice agent test set after choosing a vendor?

Yes, and you should. A shared test set becomes a regression suite: rerun it after any prompt change, model swap, or vendor update, and a drop from the previous run flags a regression before callers find it. It also gives you a fair basis for A/B testing your own versions. Reusing the set is most of its value, not a bonus.

How do you keep a voice agent comparison fair?

Hold everything constant except the agent under test: the same scenarios, the same caller profiles, the same order, and the same scoring rules against pre-written outcomes. Measure on your own data, run at production concurrency, and keep the result reproducible. Any variable you let differ between agents becomes a confound, so a fair comparison is mostly a matter of disciplined sameness.

What's the difference between a test set and a benchmark?

A test set is the fixed collection of scenarios, profiles, and expected outcomes you run. A benchmark is the act of running it and producing comparable numbers, with attention to sample size and statistics. The test set is the reusable artifact; the benchmark is the measurement you perform with it. You build the set once, then benchmark every agent and version against it over time.

Related Articles