Accent Robustness for Voice Agents: Finding and Fixing Speech Recognition Accent Bias on US Calls

On this page
Your voice agent works in the demo. It books the appointment, reads back the confirmation number and hangs up cleanly. Then it goes live on a US phone number, and a slice of callers start hearing "Sorry, could you repeat that?" three times before they give up and press zero.
Nobody flags it as an accent problem, because nobody looks at the data that way. The dashboard shows one blended word error rate, and the callers who struggle most rarely complain. They just stop calling.
This guide is for the engineer who owns that agent: what research measured about accent gaps in speech-to-text (STT), why 8 kHz telephony makes them worse, how to build and size an accent-stratified test set, how to monitor without collecting protected attributes, and which fixes work. Every number comes from a linked source or is labeled as a worked example.
A note on framing. Everyone has an accent. "Accent robustness" here means the agent works equally well for the full range of English your callers speak: regional US varieties, African American English, and English spoken by people whose first language is Spanish, Vietnamese, Hindi, Tagalog or anything else. The goal is consistent service, and the method is measurement, not guessing.
What the research measured, and what it means for your agent
Most write-ups on speech recognition accent bias cite one headline number. The useful material is in the methods sections of these six studies.
Koenecke et al. (PNAS 2020): the gap lives in the acoustic model, and in the tail
Racial disparities in automated speech recognition tested Amazon, Apple, Google, IBM and Microsoft recognizers on 19.8 hours of sociolinguistic interviews: 2,141 snippets from 73 Black speakers (from the CORAAL corpus) matched to 2,141 snippets from 42 white speakers (from Voices of California) on age, gender and snippet length. Average WER was 0.35 for Black speakers and 0.19 for white speakers. The best system, Microsoft's, scored 0.27 versus 0.15. Every system showed error rates for Black speakers nearly twice as large.
Three details matter more than the headline.
The tail is worse than the mean. More than 20% of snippets from Black speakers had WER of at least 0.5. Fewer than 2% of snippets from white speakers crossed that line (the figure caption gives 23% versus 1.6%). If WER above 0.5 means a transcript is unusable, ten times as many turns fail outright. For a voice agent, those are the turns that trigger repeat requests and hang-ups. Track the share of turns above a failure threshold, not just the average.
The language model was not the cause. The authors checked vocabulary coverage (98 to 99% for both groups) and perplexity under GPT-2 and other language models (lower for the Black speakers' transcripts, meaning easier to predict). Then they compared 206 identical short phrases spoken by both groups. The gap held: for Microsoft, 0.13 versus 0.07 on the same words. The error came from pronunciation and prosody: rhythm, pitch, vowel duration. Adding words to a vocabulary list will not close a gap that sits in the acoustic model.
Error tracked dialect density. WER rose with the density of African American English features, and the Rochester site, with the lowest density, scored close to the California sites. Speech varies within any group, and your test set needs that range.
The authors released their data and code, including the transcript normalization rules. Reuse those rules. Filler words, number formats and spelling variants can swing WER by points if two systems are normalized differently.
EdAcc (ICASSP 2023): clean benchmarks hide conversational accent gaps
The Edinburgh International Accents of English Corpus has almost 40 hours of video-call conversations between friends, covering many first- and second-language varieties of English. The best model tested, trained on 680,000 hours of data, averaged 19.7% WER on EdAcc, against 2.7% on clean read US English. All models dropped on Indian, Jamaican and Nigerian English speakers.
What it means for you: a vendor's 3% WER on a read-speech benchmark tells you almost nothing about conversational speech across accents. Your callers are talking, not reading.
WildASR (2026): accent gaps shrink on read speech, and robustness does not transfer
Back to Basics: Revisiting ASR in the Age of Voice Agents tested seven systems, including Deepgram Nova 2, GPT-4o Transcribe, Gemini, ElevenLabs Scribe V1 and Whisper Large V3. On its English accent subset (read, non-native speech from the GLOBE corpus), WER stayed in low single digits: 2.2% to 6.8% across systems. That looks reassuring until you read the rest of the paper.
On the same benchmark, simulated G.711 phone audio raised average English WER on FLEURS read speech from 4.1% to 10.4%. Robustness in one condition did not predict robustness in another. Under truncated audio, models "hallucinate plausible but unspoken content", which the authors flag as a direct risk for agents making API calls. In a reverberation sweep, they also found the 90th-percentile WER climbed faster than the mean as the distortion got worse.
What it means for you: accent results from clean, read speech are an upper bound. Test accents and the phone channel together, and watch the P90.
L2-ARCTIC study (2025): read versus spontaneous changes the ranking
A 2025 evaluation ran five ASR systems on the L2-ARCTIC corpus: 24 speakers whose first languages are Arabic, Chinese, Hindi, Korean, Spanish and Vietnamese. On read sentences, Whisper and AssemblyAI led with match error rates of 0.054 and 0.056. On spontaneous narratives, Rev.ai led at 0.063. Systems also differed widely in how they handled fillers, repetitions and self-corrections.
What it means for you: the winner on read speech was not the winner on spontaneous speech. Run your bake-off on speech that sounds like your calls.
tau-Voice (2026): accents cost more task success than noise, and it depends on the provider
tau-Voice benchmarks full-duplex voice agents from Google, OpenAI and xAI on grounded, multi-turn tasks with tool calls. In the Retail domain ablation, adding accents (through TTS voice personas) dropped pass@1 by 1 point for Google (45% to 44%), 11 points for OpenAI (71% to 60%) and 18 points for xAI (48% to 30%). Averaged over the three, accents cost 10 points, more than noise (4 points) or turn-taking (7 points). The paper's own example of the failure: the agent mishears a name, authentication fails, and the question becomes whether the agent asks the caller to spell it and then fixes the tool call.
What it means for you: accent robustness is a property of your specific stack, and it shows up as task failure, not just WER. The authors caution that TTS accents make these results "indicative rather than definitive", which leads to the next study.
Lau et al. (ISSTA 2023): synthetic accents raise false alarms
Synthesizing Speech Test Cases with Text-to-Speech? checked what happens when ASR failures found with TTS audio are replayed with human audio of the same text. Between 21% and 34% of TTS-found failures were false alarms: the human recording transcribed correctly. False alarm rates also varied by TTS engine, from 17% to 32%.
What it means for you: TTS accent personas are good for breadth and regression, but at least one failure in five may not be real. Confirm with human audio before you call something a disparity.
Why accents hurt more on 8 kHz phone calls
Most US voice agents receive caller audio as 8 kHz G.711 mu-law. Twilio Media Streams deliver it that way, and most SIP trunks negotiate it.
The phone channel deletes evidence the recognizer needs
At 8 kHz sampling, nothing above 4 kHz survives, and the classic telephone band runs roughly 300 to 3,400 Hz. Vowel identity depends mostly on the first two formants, which sit below 3 kHz, so vowels mostly survive. Many consonant contrasts do not. The energy that separates "s" from "f" from "th", or marks a final "t" or "d", sits partly or mostly above 4 kHz.
A recognizer handles a missing cue by leaning on other cues and on context. Accents change exactly those other cues: vowel quality, stress, which final consonants are released, rhythm. When the accent shifts the vowel and the channel strips the consonant, a model has less to go on than for a speaker who matches its training data.
Upsampling does not help. Resampling 8 kHz audio to 16 kHz adds samples, not information; the upper half of the spectrum stays empty. Mobile legs add another codec (AMR or similar) before the call reaches the PSTN, and packet loss concealment smears short sounds. WildASR's simulated GSM condition raised error over clean audio in all four languages it tested.
Rhythm meets the endpointer
The second mechanism isn't acoustic. Callers differ in speech rate and pause patterns. Someone speaking a second language may pause mid-sentence to find a word, and anyone reading a reference number tends to read it in chunks.
A fixed silence timer tuned on the team's own voices will end the turn in the middle of "My account number is... 4 4 7..." The STT finalizes a truncated utterance, which is where WildASR found recognizers hallucinate fluent completions. The LLM gets a confident, wrong partial, and the tool call fires with it. So an accent-related failure can start with an endpointing setting, not the recognizer. Test the whole pipeline.

Which errors matter for an agent, and which don't
WER counts every substitution, deletion and insertion equally. A voice agent doesn't. "Um, I'd like to, uh, reschedule" transcribed as "I like to reschedule" scores three errors and changes nothing. "Okafor" transcribed as "okay for" scores two errors and fails authentication.
So accent work has to move past WER. For the mechanics, see word error rate for voice agents and STT entity accuracy. Here is the taxonomy that matters.
| Error class | Example | WER impact | Agent impact | Signal to watch |
|---|---|---|---|---|
| Fillers and function words | "uh", "to", "the" dropped | Moderate | None | Ignore after normalization |
| Personal names | "Nguyen" as "win", "Okafor" as "okay for" | Low (1-2 words) | Failed lookup or auth | Lookup miss rate per slot |
| Street and place names | "Tchoupitoulas" as "chop it to less" | Low | Wrong address, failed delivery | Address validation failures |
| Alphanumerics | "4B" as "for be", "M" as "N" | Low | Wrong account, wrong unit | Checksum or format failures |
| Numbers and dates | "fifteen" as "fifty", "thirteenth" as "thirtieth" | Very low | Wrong amount, wrong appointment | Readback rejections |
| Polarity | "can" vs "can't", "do" vs "don't" | Very low | Opposite action | Contradiction in next turn |
| Intent keywords | "cancel" as "council" | Low | Wrong flow | Intent fallback rate |
| Yes and no | "yeah" missed, "nah" as "now" | Very low | Stalled confirmation | Repeat confirmations |
The rows that matter most barely move WER. Two groups at 8% and 6% WER look close to parity, but if the extra errors land on names and numbers, the entity error gap can be several times larger.
The downstream chain is predictable:
1. Wrong tool arguments. A misheard name goes straight into `lookup_customer(last_name=...)`. Our guide to tool argument accuracy covers how to score this.
2. Repeat requests. The agent asks again. Each repeat adds a full turn of latency plus the caller's time.
3. Escalations. After two or three failures, the call goes to a human, or the caller hangs up.
4. Longer handle time. Even when the task succeeds, the group that needed extra turns waited longer.
For a collections line, a clinic scheduler or a utility, a service that works worse for some callers is a fairness and compliance question, not just a CX metric.
Building an accent-stratified test set
Your test set needs three properties. It must cover the speech varieties your callers use. It must sound like your phone channel. And it must contain your entities, because generic sentences don't exercise name and address recognition.
Public corpora: what's available and what the licenses allow
| Corpus | What it contains | License | Commercial testing | Best use |
|---|---|---|---|---|
| CORAAL | Sociolinguistic interviews, African American English, several US sites | CC BY-NC-SA 4.0; some components available under additional licenses | Contact maintainers | Conversational AAE, the corpus behind Koenecke et al. |
| L2-ARCTIC | 24 speakers, L1 Arabic, Chinese, Hindi, Korean, Spanish, Vietnamese; read plus narrative | CC BY-NC 4.0 | Contact maintainers for other uses | Second-language English, controlled sentences |
| Speech Accent Archive | Native and non-native speakers reading the same paragraph | CC BY-NC-SA 4.0 | Non-commercial only | Same text across many accents |
| Common Voice | Crowd-recorded read sentences with self-reported accent field | CC0 | Yes, with a no re-identification term | Broad accent coverage, filter by accent |
| EdAcc | ~40 h conversations, many L1 and L2 varieties | CC BY-SA | Share-alike applies | Conversational, international varieties |
NC licenses. Testing a product you sell is hard to call non-commercial. If you plan to use CORAAL, L2-ARCTIC or the Speech Accent Archive inside a company, get a legal read or ask the maintainers. CORAAL offers additional licenses for some components, and the L2-ARCTIC page invites requests for uses the NC license doesn't cover.
Speaker concentration in Common Voice subsets. Common Voice accent labels are self-reported and multi-select. Curated accent subsets can be tiny. The Southern American English subset of Common Voice 26.0 on Mozilla Data Collective has 1,175 clips from 25 speakers, 1.79 hours, and two speakers account for 78% of the clips. Its dev and test splits are each one speaker. A score on that split is a score for one person. Count speakers, not clips.
Recording your own callers, with consent
Public corpora never contain your entities: your customers' surnames, your street names, your SKU formats. The strongest test set is a consented panel that works through your scenarios over a real phone line.
- Recruit participants who describe their own speech variety, in their own words, and consent in writing to recording and evaluation use. Store the self-description as a test-set label. Never infer it.
- Give each participant 10 to 20 scripted entity prompts (a name, an address, a date, an account number) and 3 to 5 open tasks ("call to move your appointment to next week").
- Have them call your staging number from their own phones. That captures real carriers, codecs and rooms.
- Pay participants, set a retention period, and keep recordings out of training data unless the consent covers it. If recordings include real PII, apply the same controls you use for production calls; see PII leakage detection.
TTS-synthesized accents: useful, with limits
TTS accent voices let you generate thousands of test calls overnight, which makes synthetic callers the right tool for coverage and regression.
They are the wrong tool for measuring a disparity. The ISSTA 2023 result says 21 to 34% of TTS-found failures were false alarms. TTS accents are also a model's idea of an accent, usually a stereotyped, over-regular one. A rule that works: TTS findings open a ticket, human audio confirms it. A disparity number goes in a report only when it comes from human speech.
Make every clip sound like your phone line
Public corpora are mostly wideband, a channel your callers never use. Convert every clip to your production format before scoring:
# Illustrative: convert a corpus clip to 8 kHz mono G.711 mu-law, like a Twilio leg
ffmpeg -i clip.wav -ar 8000 -ac 1 -c:a pcm_mulaw clip_8k_mulaw.wavThen send it with the same encoding parameters as production (for Deepgram, `encoding=mulaw&sample_rate=8000`). Better still, play a subset through a real call loop so carrier processing is included. For noise conditions, layer the methods from testing STT with background noise.
The test matrix
Cross speech varieties with channel conditions and entity types. Not every cell needs equal depth, but every cell needs some.

A starting matrix for a US support agent might have six groups: a reference group matching the variety the agent was tuned on, two regional US varieties, African American English, and two or three second-language groups that match your customers. Choose groups from your own caller mix, not a generic list.
Name groups by speech variety, using labels participants chose, not by race or ethnicity. "Spanish-L1 English" describes speech. A race label describes a person and tells you nothing about how they talk.
Callers who switch languages mid-call need a separate design; see multilingual accuracy for voice agents.
Metrics per group, the disparity ratio and the CI math
Define the metrics before you run anything
Score each group on these, with fixed definitions:
- WER, after one normalization applied to every system: lowercase, strip punctuation, remove fillers, expand symbols. Reuse the Koenecke rules as a baseline.
- Entity error rate (EER): entities where the extracted value doesn't match the reference after entity-level normalization, divided by total entities. Score this on what reaches the tool call, not on the raw transcript.
- Task success: the task completed with correct arguments, judged against ground truth.
- Repeat-request rate: agent turns asking the caller to repeat or rephrase, per call.
- P90 utterance WER: the tail, per Koenecke and WildASR.
The disparity ratio
For an error metric, the disparity ratio for group g is:
DR(g) = EER(g) / EER(reference)
A DR of 1.0 is parity. Koenecke's averages give a WER ratio of 0.35 / 0.19 = 1.84. For success metrics, report the gap in percentage points: success(reference) minus success(g).
Report the worst group, not the average across groups. A blended number lets one well-served majority group hide a badly served minority group.
Here is an illustrative gate. These thresholds are our suggested starting points, not industry standards. Set your own based on what each failure costs.
| Metric | Definition | Pass | Investigate | Block release |
|---|---|---|---|---|
| EER disparity ratio | EER(g) / EER(ref), worst group | CI upper bound under 1.25 | Point estimate 1.25-1.5 | CI lower bound above 1.5 |
| Task success gap | success(ref) - success(g) | Under 3 pts | 3-8 pts | Over 8 pts with CI excluding 0 |
| Repeat-request rate | Repeats per call, ratio to ref | Under 1.3x | 1.3-2x | Over 2x |
| P90 utterance WER | 90th percentile per group | Under 0.3 | 0.3-0.5 | Over 0.5 |
Why you must resample speakers, not utterances
Utterances from the same speaker are not independent. If one participant's voice trips the recognizer, all twenty of their utterances fail together. Treat 400 utterances from 20 speakers as 400 independent samples and your confidence interval will be far too narrow. You'll report a disparity that is one person's bad day.
The standard correction is the design effect:
DE = 1 + (m - 1) x rho
where m is utterances (or entities) per speaker and rho is the intraclass correlation, how much a speaker's errors cluster. Divide your raw sample size by DE to get the effective sample size, or multiply your required sample by DE when planning.
Worked example: how much data per group?
Assume a reference EER of 10% and you want to detect a group at 15% (a disparity ratio of 1.5), with two-sided alpha of 0.05 and 80% power. The two-proportion formula:
n = (z_alpha/2 + z_beta)^2 x [p1(1 - p1) + p2(1 - p2)] / (p2 - p1)^2
n = (1.96 + 0.84)^2 x [0.10 x 0.90 + 0.15 x 0.85] / 0.05^2
n = 7.85 x 0.2175 / 0.0025 = about 683 entities per group
That assumes independence. With 10 entities per speaker and rho = 0.1 (an illustrative assumption; estimate yours from pilot data), DE = 1 + 9 x 0.1 = 1.9, so you need about 1,297 entities per group, or about 130 speakers at 10 entities each.
For task success, detecting a drop from 80% to 70% needs about 290 tasks per group before clustering.

Two practical conclusions. First, small gaps are expensive to detect. A 3-point EER gap needs thousands of entities per group. Decide what gap matters before you recruit. Second, adding speakers buys more than adding utterances per speaker. Ten utterances from each of 100 people beats fifty from each of 20.
Code: per-group WER and EER with speaker-level bootstrap CIs
This script takes one CSV row per utterance and reports WER, EER and the EER disparity ratio per group, with 95% confidence intervals from a speaker-level bootstrap. It uses jiwer for edit counts. Simplified: adapt the normalizer and entity matching to your data.
"""Per-group WER and entity error rate with speaker-level bootstrap CIs.
Simplified: adapt normalize() and normalize_entity() to your data.
CSV columns: speaker_id, group, reference, hypothesis, ref_entities, hyp_entities
(entity columns are JSON, e.g. {"last_name": "Okafor", "unit": "4B"})"""
import csv, json, re, random, sys
from collections import defaultdict
import jiwer # pip install jiwer
FILLERS = {"um", "uh", "mm", "hm", "mhm", "huh"}
def normalize(text):
text = text.lower().replace("$", " dollar ")
text = re.sub(r"[^a-z0-9' ]+", " ", text)
return " ".join(w for w in text.split() if w not in FILLERS)
def normalize_entity(value):
# "4B" == "4 b", "O'Neil" == "oneil"
return re.sub(r"[^a-z0-9]", "", value.lower())
def speaker_stats(path):
stats = defaultdict(lambda: {"group": None, "errors": 0, "words": 0,
"ent_wrong": 0, "ent_total": 0})
with open(path, newline="") as f:
for r in csv.DictReader(f):
s = stats[r["speaker_id"]]
s["group"] = r["group"]
ref, hyp = normalize(r["reference"]), normalize(r["hypothesis"])
if ref:
o = jiwer.process_words(ref, hyp)
s["errors"] += o.substitutions + o.deletions + o.insertions
s["words"] += o.substitutions + o.deletions + o.hits
ref_ents = json.loads(r["ref_entities"] or "{}")
hyp_ents = json.loads(r["hyp_entities"] or "{}")
for slot, val in ref_ents.items():
s["ent_total"] += 1
if normalize_entity(hyp_ents.get(slot, "")) != normalize_entity(val):
s["ent_wrong"] += 1
return stats
def rates(speakers):
words = sum(s["words"] for s in speakers)
ents = sum(s["ent_total"] for s in speakers)
wer = sum(s["errors"] for s in speakers) / words if words else float("nan")
eer = sum(s["ent_wrong"] for s in speakers) / ents if ents else float("nan")
return wer, eer
def ci(values):
v = sorted(x for x in values if x == x) # drop NaN
return round(v[int(0.025 * len(v))], 3), round(v[int(0.975 * len(v)) - 1], 3)
def report(by_group, reference, n_boot=2000, seed=7):
rng = random.Random(seed)
ref_spk = by_group[reference]
_, ref_eer = rates(ref_spk)
for g, spk in sorted(by_group.items()):
wer, eer = rates(spk)
draws = []
for _ in range(n_boot):
a = [rng.choice(spk) for _ in spk] # resample speakers
b = [rng.choice(ref_spk) for _ in ref_spk]
(wa, ea), (_, eb) = rates(a), rates(b)
draws.append((wa, ea, ea / eb if eb else float("nan")))
row = {"speakers": len(spk),
"entities": sum(s["ent_total"] for s in spk),
"wer": round(wer, 3), "wer_ci": ci(d[0] for d in draws),
"eer": round(eer, 3), "eer_ci": ci(d[1] for d in draws)}
if g != reference and ref_eer:
row["eer_ratio"] = round(eer / ref_eer, 2)
row["eer_ratio_ci"] = ci(d[2] for d in draws)
print(g, row)
if __name__ == "__main__":
stats = speaker_stats(sys.argv[1])
groups = defaultdict(list)
for s in stats.values():
groups[s["group"]].append(s)
report(groups, reference=sys.argv[2])Run it as `python acc_eval.py results.csv reference_group`. On a toy file with 12 speakers per group, it printed an EER ratio of 2.2 with a 95% interval of roughly 1.2 to 4.4. With 12 speakers, a ratio of 2.2 fits anything from a mild gap to a severe one. Recruit more speakers before you draw conclusions, and keep the normalizer under version control: a changed normalizer is a changed metric.
Monitoring accent gaps in production without collecting protected attributes
Pre-launch testing tells you where you started. Models update, prompts change, and caller mix shifts. You need a production signal, and you need it without labeling callers by race, ethnicity or national origin.
Don't build an accent classifier on your callers
It's tempting to run an accent-identification model on live calls and slice metrics by its output. Don't. Inferring accent from voice is close to inferring national origin or ethnicity. It creates a sensitive dataset you didn't need, mislabels many people and raises compliance questions. You can find and fix gaps without it.
Use proxy signals that measure the failure, not the person
These measure whether the agent understood, not who the caller is:
- Repeat-request rate: agent turns asking for repetition, per call and per slot.
- Slot re-prompt count: how many attempts each entity took (name, address, date of birth).
- Readback rejection rate: how often callers say "no" to a confirmation.
- Modality fallback rate: how often a slot ends in DTMF, SMS link or spelling mode.
- Escalation after slot failure: transfers that follow two or more failed attempts on one slot.
- Low-confidence turn share: turns with STT word confidence under a threshold, from your provider's per-word confidence.
Slice these by non-sensitive operational dimensions: phone line or DID, campaign, language setting, carrier and codec, time of day, and broad region from the caller's area code. Area code is coarse, and that's the point. It's enough to show, for example, that calls into one regional line need far more repeats than calls into another, without saying anything about any individual caller. Use these slices to find where to look, never to make decisions about individual callers.
Log the raw events these metrics need on every call; what to log on every voice agent call has the schema.
Close the loop with opt-in data
- A post-call survey with one question: "Did the assistant understand you?" Add an optional, self-described "How would you describe your accent or the variety of English you speak?" as a free-text field. Make it clearly optional, explain why you ask, and store it separately from account data.
- A consented recording panel: when a segment's proxy signals drift, pull calls only from callers who opted in to quality review, and have a human label what went wrong.
A workable alert rule (an illustrative starting point): flag any segment with at least 300 calls in a rolling two weeks whose repeat-request rate exceeds 1.5 times the fleet rate, then pull 30 consented calls from that segment for human review.
Fixes that work, in order of cost
1. Configuration: the right model and locale per line
Robustness is provider-specific. tau-Voice measured a 1-point accent drop for one provider and 18 points for another. Run your stratified set against two or three candidates before committing. The bake-off method is in our STT evaluation guide.
Check what locale options your provider offers. Per Deepgram's models and languages page, Nova-3 accepts English locale codes `en-US`, `en-AU`, `en-GB`, `en-IN` and `en-NZ`, while Flux has one English model, `flux-general-en`, described as "English (all accents)". Whether a locale hint helps your callers is an empirical question. A US caller who grew up speaking Hindi is not necessarily better served by `en-IN`. Test it on your panel. If you are choosing between those two Deepgram models, Deepgram Flux vs Nova-3 covers the trade-offs.
You can't switch models per caller without guessing their speech variety, and you shouldn't guess. Set models per line or campaign when the caller base is known, or run a second recognizer pass on a slot that failed twice.
2. Keyterm prompting: bias the recognizer toward the right answer
Keyterm prompting won't fix an acoustic gap across the board, as Koenecke showed. It does work well for the entity class that hurts most: a known, finite set of names and places. If the CRM says the caller on this number is "Okafor", tell the recognizer before asking for the last name.
Deepgram's keyterm docs list the details that matter:
- `keyterm` works on Nova-3 and Flux, up to 500 tokens per request, and Deepgram recommends focusing on 20 to 50 terms.
- Repeat the parameter for each term: `keyterm=Okafor&keyterm=Tchoupitoulas`. Comma-separated or weighted values don't error. They are accepted as one literal term and silently boost nothing.
- You can replace keyterms mid-stream with a `Configure` message carrying a `keyterms` array, on both Nova-3 and Flux. Load the caller's name before the name slot, your service area's street names before the address slot, then clear them.
- Mid-stream Nova-3 keyterm updates work on the global endpoint but, as of this writing, not on the EU, Australia or India regional endpoints.
{"type": "Configure", "keyterms": ["Okafor", "Tchoupitoulas Street", "Unit 4B"]}On OpenAI, the speech-to-text guide documents `prompt`, `keywords` and `languages` for `gpt-transcribe`. For `whisper-1`, prompts are limited to 224 tokens. If you use Whisper, test prompts on your accent panel, because a prompt that helps one group can shift errors onto another.
3. Confirmation strategies sized to the risk of each entity
Design the conversation so a misrecognition gets caught before it reaches a tool call.
| Entity | Risk if wrong | Confirmation strategy | Fallback after 2 failures |
|---|---|---|---|
| Last name | Failed lookup, wrong account | Read back, then ask to spell if lookup fails | Spell letter by letter, with "B as in boy" readback |
| Account or member number | Wrong account | Read back in chunks of 3 to 4 digits | DTMF keypad entry |
| Date of birth | Failed verification | Read back as "March fifteenth, nineteen eighty" | DTMF as MMDDYYYY |
| Street address | Wrong delivery | Validate against an address API, then read back | SMS a form link |
| Dollar amount | Wrong payment | Read back with "dollars and cents" | DTMF entry |
| Yes or no on a consequential action | Wrong action | Restate the action, ask again | Transfer to a human |
Two rules make this work. Switch modality instead of asking the same question a third time; a third identical request rarely works and feels like blame. And never make the caller start over after a fallback. DTMF has its own failure modes, so test it like any other path; see testing DTMF navigation.
4. Endpointing tuned for different speech rhythms
If callers get cut off mid-entity, fix the turn detector before the recognizer. In LiveKit Agents, endpointing lives under `turn_handling`. Per the turn handling reference, `endpointing` takes `mode` (`"fixed"` or `"dynamic"`), `min_delay`, `max_delay` and `alpha`. Dynamic mode adapts the delay within the min and max range from each session's own pause statistics, which suits callers whose rhythm differs from the default.
You can also widen the window just for entity slots with `session.update_options(endpointing_opts={...})`, which takes effect on the next turn:
# Illustrative, LiveKit Agents (Python): give callers more time while reading a number
session.update_options(
endpointing_opts={"mode": "fixed", "min_delay": 1.2, "max_delay": 4.0},
)
# ...after the slot is captured, return to adaptive endpointing
session.update_options(
endpointing_opts={"mode": "dynamic", "min_delay": 0.5, "max_delay": 3.0, "alpha": 0.9},
)On Deepgram Flux, the equivalent knobs are `eot_threshold` and `eot_timeout_ms`. Test any change on every group: a longer delay helps a caller who pauses mid-number and adds dead air for one who doesn't. See best endpointing for voice agents and disfluency robustness for the details.
5. Escalation rules that respect the caller's time
Escalation is part of accent robustness, not a failure. A good rule set:
- After two failed attempts on one slot in the same modality, switch modality.
- After a failed modality switch, or three failed slots in one call, offer a human.
- Pass the confirmed fields and the failed slot to the human, so the caller doesn't repeat everything.
- Track escalation-after-slot-failure per segment. It's one of your production proxy signals.
Escalation accuracy and handoffs covers how to test that handoffs carry context.
How to test accent robustness in your voice agent
This is the protocol end to end. Plan on two to four weeks for the first run, then hours for each rerun.
1. Pick groups from your caller base. Use call volume by line and region, plus any opt-in survey data, to choose four to seven speech-variety groups, including a reference group. Write down the smallest gap you care about detecting.
2. Size the sample. Use the two-proportion formula and a design effect from pilot data. As a floor, plan 30 or more speakers per group, with 10 or more entity utterances each. Plan more if the gap you care about is small.
3. Assemble audio. Combine consented panel calls over real phone lines with licensed public corpora (check NC terms) and TTS accent personas for breadth. Tag every clip with its source.
4. Match the channel. Convert corpus audio to your production format (8 kHz mu-law for most US telephony), and run a subset through a real call loop.
5. Script your entities. Use the names, streets, number formats and dates your agent actually collects. Build a golden dataset with reference transcripts and reference entity values.
6. Run the full pipeline, not just STT. Send calls through your agent so endpointing, keyterms, confirmation and tool calls are all exercised. Capture transcripts, extracted entities, tool arguments and task outcomes. Tool-calling test cases shows how to assert on arguments.
7. Score per group. Run the bootstrap script for WER and EER, and compute task success gap, repeat-request rate and P90 WER per group. Apply your gate.
8. Confirm TTS findings with humans. Any gap found only in TTS audio is a hypothesis. Reproduce it with human recordings before reporting it.
9. Fix, then rerun the whole matrix. Apply fixes in cost order. A fix for one group can shift errors to another, so rerun every group.
10. Wire production monitoring. Ship the proxy metrics, segment slices, opt-in survey and alert rule. Rerun the stratified suite on every STT, prompt or model change.
Run this yourself when you have the panel and the time. An independent evaluator helps when you lack a stratified panel and phone-path harness, when you want an STT bake-off nobody with a stake scored, or when a regulator or enterprise customer asks for evidence. Evalgent does that work as a third party; see independent voice AI evaluation.
Frequently asked questions
What is speech recognition accent bias?
Speech recognition accent bias is a systematic difference in transcription accuracy between groups of speakers with different speech varieties. Koenecke et al. found five commercial recognizers averaged 0.35 WER for Black speakers versus 0.19 for white speakers, with the gap traced to acoustic models. In voice agents it shows up as failed name lookups, repeat requests and escalations.
How do I test a voice agent with different accents?
Build a stratified test set with speech-variety groups drawn from your caller base, ideally 30 or more speakers per group. Mix consented recordings over real phone lines, licensed corpora and TTS voices for breadth. Run calls through the full agent, then score WER, entity error rate and task success per group with speaker-level confidence intervals.
Why are accent errors worse on phone calls?
US telephony usually carries 8 kHz G.711 audio, which removes everything above 4 kHz, including cues that separate many consonants. Accents shift other cues like vowels and stress, so the recognizer has less redundant evidence. Fixed endpointing can also cut off callers who pause mid-sentence, and truncated audio is where recognizers tend to hallucinate.
Are TTS-generated accents good enough for STT testing?
They are good for coverage and regression, not for measuring disparities. An ISSTA 2023 study found 21 to 34% of ASR failures found with TTS audio were false alarms that human recordings of the same text didn't reproduce. The tau-Voice authors also call their TTS-accent results indicative. Confirm every TTS-found gap with human speech.
How many test utterances do I need per accent group?
It depends on the gap you want to detect. Detecting an entity error rate of 15% against 10% at 80% power needs about 683 independent entities per group. Because errors cluster by speaker, multiply by a design effect: with 10 entities per speaker and intraclass correlation of 0.1, that's about 1,300 entities, roughly 130 speakers.
Does Deepgram handle different English accents?
Deepgram Nova-3 accepts English locale codes en-US, en-AU, en-GB, en-IN and en-NZ, and its Flux model has one English option described as all accents. Keyterm prompting works on both and can be updated mid-stream. How well either handles your callers is something to measure on your own stratified, phone-quality test set.
How can I monitor accent gaps without collecting demographic data?
Track behavioral signals that measure misunderstanding: repeat-request rate, slot re-prompts, readback rejections, DTMF fallbacks and escalations after slot failure. Slice them by line, campaign, carrier and broad area-code region. Add an opt-in post-call survey and a consented review panel. Don't run accent classifiers on callers.
What is the fastest fix for accent-related STT errors?
For names and addresses, load expected values as keyterms right before the slot, then confirm with readback and switch to spelling or DTMF after two failures. If callers get cut off mid-entity, widen endpointing for that slot. These changes need no model swap and target the errors that break tool calls.
The bottom line
Speech recognition accent bias in a voice agent shows up as entity errors, repeat requests and escalations that a blended WER hides, so measure it per speech-variety group, on phone-quality audio, with speaker-level confidence intervals. Fix it in cost order: configuration and keyterms, confirmation and fallbacks, endpointing, then escalation rules, and rerun the full matrix after every change.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more