Evalgent
Back to Blog
Voice AI Testing

How to A/B test voice variations in Deepgram

Deepesh Jayal
10 min read
How to A/B test voice variations in Deepgram

Teams pick their voice agent's voice the same way they pick a ringtone: someone listens to a few options and chooses the one they like. That is not a test — it is a preference. The voice a caller hears shapes whether they understand the agent, trust it, and complete their task, and the one that sounds nicest to you in a quiet room is not always the one that performs best on real calls. A/B testing voices in Deepgram replaces the vibe check with evidence. Evalgent runs exactly this kind of comparison, and this guide shows how.

This is specifically about testing the voice — a different job from A/B testing agent versions or prompts, which our A/B testing voice agents guide covers. Here, everything else stays fixed and only the voice changes.

Voice A/B test: a controlled comparison where two or more text-to-speech voices are run against the same scenarios, with all else held constant, to measure which voice produces better caller outcomes.

Why the voice deserves its own A/B test

It is tempting to treat the voice as cosmetic, but it sits directly in the caller's experience. A voice that is hard to understand raises re-ask rates. A voice that mispronounces names or numbers erodes trust on exactly the details that matter. A slower voice adds latency to every turn. None of these show up when you audition voices by ear on a demo script.

The reason a formal A/B test matters is that voice effects are real but small and easy to misjudge. A voice can sound warmer and more pleasant while actually lowering task completion, because pleasantness and intelligibility are not the same thing. The only way to know which voice serves your callers is to measure outcomes, not impressions, across realistic calls.

How voices work in Deepgram

To A/B test voices in Deepgram, you need to know what a "voice variation" actually is on the platform. Deepgram's text-to-speech, Aura, offers a catalog of named voices. As the Deepgram TTS models docs describe, each is identified in the form `[model]-[voice]-[language]` — for example, `aura-2-thalia-en`. Switching that string is how you switch voices.

There are also two model families to consider. Per Deepgram, Flux is the newer, streaming-first family built voice-agent-first with cross-turn voice consistency, while Aura offers the broadest voice catalog across English and Spanish. Within the Voice Agent API, you set your text-to-speech model in the Settings Message, as the voice agent TTS docs explain. That single setting is your A/B lever: change only the voice model, keep the prompt, LLM, and everything else identical, and you have a clean comparison.

What counts as a voice variation

A voice variation is anything you can change on the text-to-speech side while holding the rest of the agent constant. In Deepgram, the practical variations are a different named Aura voice, a different model family such as Flux versus Aura, or a language-specific voice for a multilingual line. Each of these is a candidate to test.

The discipline is to change one thing at a time. If you swap the voice and the prompt together, a difference in outcomes tells you nothing about which change caused it. A voice A/B test isolates the voice so the result is attributable.

What to measure when A/B testing voices

The winning voice is the one that produces better caller outcomes, so measure outcomes rather than aesthetics. Track these across each voice.

MetricWhat it reveals about the voice
Comprehension / re-ask rateHow often callers ask the agent to repeat
Task completionWhether callers reach their goal
Naturalness (MOS)Perceived human-likeness and clarity
Caller sentimentFrustration or ease during the call
Latency / TTFB per voiceAdded delay each voice introduces
Pronunciation accuracyHow each voice handles names and numbers

Naturalness matters, but on its own it is a trap — it measures how the voice sounds, not whether callers succeed. Read it alongside task completion and re-ask rate, so a pleasant voice that quietly lowers completion cannot win by charm. Pronunciation deserves its own attention, since voices differ on exactly the load-bearing words, a point our pronunciation guide covers directly.

How to A/B test voices in Deepgram

The method is a clean, controlled comparison with outcomes as the verdict.

1. Pick candidate voices — Shortlist two or three Aura voices, or a Flux-versus-Aura pairing, worth comparing for your use case.

2. Change only the voice — Swap the text-to-speech model in the Settings Message and hold the prompt, LLM, and audio path identical.

3. Run the same scenarios — Drive each voice through the same set of realistic calls, or split live traffic evenly between them.

4. Measure outcomes — Record comprehension, task completion, sentiment, and per-voice latency for each variant, not just how each sounds.

5. Check significance — Gather enough calls per voice that the difference is real, not noise from a small sample.

6. Decide and re-test — Choose the voice that wins on outcomes, and re-run the comparison when your prompt, model, or voice catalog changes.

Live split test vs synthetic comparison

There are two honest ways to run this. A live split test routes a share of real callers to each voice and compares outcomes — accurate, but slow, and it exposes real callers to the losing voice while you learn. A synthetic comparison runs the same voices through automated calls before launch, so you choose the better voice without spending live traffic to find out.

Most teams do both: a synthetic comparison to pick the launch voice, then a smaller live split to confirm the result holds with real callers. The synthetic pass is where you can test many voices cheaply and across accents and noise, which a live split cannot do quickly. This pre-launch, outcome-based comparison is the heart of what Evalgent does.

Common mistakes

A few habits turn a voice A/B test into theater, and they are easy to avoid.

Judging by ear alone is the big one — a listening session measures your taste, not your callers' success. Changing more than the voice, like swapping the prompt at the same time, makes the result unattributable. Too small a sample lets random noise crown a winner that will not hold. Ignoring latency differences hides that a nicer voice may be a slower one. And testing on clean audio or a single accent flatters every voice equally, so you never see which one holds up for the callers who struggle most. The whole point of the test is to see the voice as your real callers hear it, a theme across our AI voice agent testing pillar.

What a winning result looks like

A voice A/B test does not always produce a dramatic winner, and that is fine. Sometimes two voices perform within noise of each other on outcomes, in which case you can pick the faster one or the one your brand prefers — the test has told you the choice is safe. What you are really watching for is a voice that quietly underperforms: similar naturalness, but a higher re-ask rate or lower completion, which means callers are struggling in ways a listening session would never reveal.

Treat the result as a decision with evidence behind it, not a leaderboard. The value of the test is not crowning a champion but ruling out the voice that would have quietly cost you completed calls in production.

A/B testing Deepgram voices with Evalgent

Evalgent compares voices on outcomes, across the callers who actually reveal the difference. Scenarios drive each Aura voice through the same realistic conversations, so the only variable is the voice. Profiles vary caller accent, pace, and line noise across the variants, since a voice that is clear for one caller can be muddy for another. Metrics record comprehension, task completion, sentiment, per-voice latency, and pronunciation accuracy, with thresholds you set, so the verdict is evidence rather than taste. Evaluations run the comparison as automated batches before launch, letting you weigh many voices cheaply. Reviews let you replay the same call on each voice and hear the difference in context.

The result is a voice choice you can defend: the Deepgram voice that measurably helps callers succeed, chosen before launch rather than by committee. For the wider Deepgram picture, see our Deepgram STT testing guide, and for the voice-quality side, the TTS evaluation guide.

Conclusion

A/B testing voices in Deepgram turns a preference into a measurement: run the same scenarios across each Aura voice, change only the voice, and let comprehension, completion, and latency decide. The voice that sounds nicest to you is not always the one that helps callers most.

Pick the voice on evidence, confirm it with real callers, and re-test when your stack changes. A voice is not a cosmetic choice — it is part of whether the call works.

Frequently asked questions

How do you A/B test voices in Deepgram?

Shortlist two or three Aura voices, then change only the text-to-speech model in the Voice Agent's Settings Message while holding the prompt, LLM, and audio path constant. Run each voice through the same realistic scenarios, or split live traffic evenly, and measure comprehension, task completion, sentiment, and latency. Choose the voice that wins on outcomes, not on how it sounds.

How do you change the voice in the Deepgram Voice Agent API?

Set the text-to-speech model in the Settings Message for your Voice Agent. Deepgram voices follow a `[model]-[voice]-[language]` format, such as `aura-2-thalia-en`, so switching that string switches the voice. Because it is a single setting, it is easy to change only the voice while keeping everything else identical, which is exactly what a clean A/B test requires.

Which Deepgram Aura voice is best for a voice agent?

There is no universal best — it depends on your callers and use case. A voice that sounds pleasant can still lower task completion, so the right choice is the one that measurably performs on your calls. Shortlist a few Aura voices, run them through the same scenarios, and let comprehension, completion, and latency decide rather than a listening preference.

What is the difference between Deepgram Aura and Flux voices?

Per Deepgram, Flux is the newer, streaming-first text-to-speech family built voice-agent-first, with features like cross-turn voice consistency, while Aura offers the broadest catalog of voices across English and Spanish. Both are valid A/B candidates; treat a Flux-versus-Aura pairing as one voice variation to test, changing only the model and measuring the outcome difference.

What should you measure when testing voices?

Measure caller outcomes, not aesthetics: comprehension or re-ask rate, task completion, naturalness or MOS, caller sentiment, per-voice latency, and pronunciation accuracy on names and numbers. Naturalness alone is misleading, because a pleasant voice can still lower completion. Reading outcome metrics alongside naturalness is what keeps a charming but less effective voice from winning the test.

Does the voice affect voice agent performance?

Yes. The voice affects how well callers understand the agent, how much they trust it, and how long each turn takes. A hard-to-understand voice raises re-ask rates, a mispronouncing voice erodes trust on key details, and a slower voice adds latency to every response. These are measurable performance effects, which is why the voice deserves a real A/B test.

How many calls do you need to A/B test a voice?

Enough that the difference between voices is statistically real rather than noise from a small sample. The exact number depends on how large the effect is and how variable your calls are — smaller differences need more calls to detect. Running the comparison with automated synthetic callers makes it practical to gather a large, consistent sample quickly before launch.

Should you A/B test voices live or before launch?

Both, in sequence. Run a synthetic comparison before launch to pick the better voice without exposing real callers to the losing one, testing many voices cheaply across accents and noise. Then run a smaller live split test to confirm the result holds with real traffic. The pre-launch pass does the heavy lifting; the live split validates it.

Related Articles