STT evaluation: how to test speech-to-text for voice agents (2026)

On this page
Speech-to-text is the first stage of every voice agent. It sets the ceiling for everything after it. If the transcript is wrong, the language model reasons over the wrong words. No prompt can fix that.
This guide is for engineers building voice agents on stacks like Pipecat, LiveKit, or Vapi. It covers the metrics, the test-set design, the normalization rules, and a working Python harness. It also covers fair provider comparisons and regression tests for vendor model updates. All example numbers are illustrative unless a source is linked.
What is STT evaluation?
STT evaluation: the practice of measuring a speech-to-text system's accuracy, latency, and robustness on audio that matches production, broken down by caller cohort.
Speech to text evaluation answers two questions. How often does the system get the words right? And under what conditions does it fail? A model can look excellent on a clean benchmark and collapse on a noisy phone line. Your job is to find that gap before callers do.
The output is never one number. "The model is 95% accurate" predicts little. "It scores 95% on clean audio and 78% on accented calls in street noise" predicts production. That second sentence is what STT evaluation produces. If you are unsure about the terms, our STT vs ASR guide explains why they mean the same thing in practice.
Why generic WER misleads for voice agents
Word error rate is the standard accuracy metric. It is necessary. It is not sufficient for voice agents, for four reasons.
First, WER weights every word equally. "The" and the third digit of an account number count the same. A voice agent fails on the digit, not on "the".
Second, WER ignores time. A streaming agent acts on partial transcripts and endpoint signals. A model can score low WER on finals yet flip partials constantly, or finalize late.
Third, WER depends on normalization. The same transcript can score 87.5% or 0% depending on text rules. We show that exact case below.
Fourth, averages hide cohorts. A 7% blended WER can contain an accent group at 25%. Those callers churn while the dashboard stays green.
Entity accuracy: the share of critical values, such as account numbers, dates, names, emails, and addresses, that the transcript captures exactly. One wrong character counts as a miss.
How do you evaluate speech to text accuracy?
You compare the system's transcript to a human-verified reference. Then you count the edits needed to turn one into the other.
WER = (S + D + I) / N
S is substitutions, D is deletions, I is insertions, and N is the number of reference words. Lower is better. Because insertions count, WER can exceed 100%. The open-source jiwer library computes WER, CER, and alignments, and it shows each substitution, insertion, and deletion.
Character error rate applies the same formula to characters. It is more sensitive to small errors in names and IDs. Our WER vs CER guide covers when each one fits.
The math is the easy part. The hard part is the reference audio. It must represent your callers, with their accents, noise, lines, and vocabulary. Without that, you measure speech recognition on the wrong distribution.
The STT evaluation metrics that matter for voice agents
A complete assessment tracks accuracy, meaning, and time together. The table lists the core set. Pass bars are illustrative starting points. Set your own from the cost of each error.
| Metric | What it catches | How to compute | Illustrative pass bar |
|---|---|---|---|
| WER | Overall transcription errors | (S + D + I) / N on normalized text | Within 2 points of your best cohort |
| CER | Small errors in names and IDs | Same formula at character level | Track trend per cohort |
| Entity accuracy | Wrong digits, names, emails, dates | Exact match of tagged values | 98% or higher on IDs |
| Keyterm recall | Missed product names and jargon | Share of expected keyterms present | 95% or higher |
| Formatting accuracy | "fifteen" vs "15", "$50" vs "fifty dollars" | Compare formatted output to spec | Matches downstream parser |
| Semantic error rate | Errors that change intent | Label or judge meaning-changing errors | Under 2% of utterances |
| Partial stability | Interim text rewriting itself | Share of partials that edit earlier words | Low and stable across releases |
| Time to first partial | Slow streaming start | First interim time minus speech start | p50 under 400 ms |
| Time to final | Slow turn completion | Final time minus speech end | p95 within your turn budget |
| Endpointing accuracy | Cut-offs and late turn ends | Compare detected end to labeled end | Few premature cut-offs |
| Hallucinated insertions | Words invented on silence or noise | Any words on silence-only clips | Near zero |
Accuracy and meaning metrics
Entity accuracy is usually the metric that decides whether a call succeeds. Tag every critical value in your references. Then check for an exact match after light normalization. Our STT entity accuracy guide goes deeper on tagging schemes.
Keyterm recall checks whether domain vocabulary survives. Product names, drug names, and plan tiers are common failures. Most vendors now accept vocabulary hints. Deepgram documents keyterm prompting. OpenAI's speech-to-text guide describes `prompt` and `keywords` parameters. Test recall with and without hints, because hints can also cause false insertions.
Formatting accuracy matters when a parser reads the transcript. If your tool call expects "15" and the STT returns "fifteen", the call fails. Test smart formatting settings as part of the model config.
Semantic error rate counts errors that change meaning. "Cancel" heard as "can sell" is one substitution. It is also a completely different intent. Label these by hand on a sample, or use a judge model with spot checks.
Streaming, latency, and endpointing metrics
Voice agents consume streams, not files. So measure the stream.
Most streaming APIs send interim results that later get corrected. Deepgram's interim results docs explain that early guesses near the audio edge are more likely wrong. Its responses mark finality with `is_final`. Its endpointing feature sets `speech_final` after a silence threshold. That threshold defaults to 10 ms and is configurable.
Endpointing now often mixes silence and semantics. AssemblyAI's turn detection docs describe a model-based trigger plus a silence fallback. Defaults include a 400 ms minimum silence when confident and a 2,400 ms maximum silence. Test the settings you will ship, not the defaults.
Deepgram's guide to measuring streaming latency makes two useful points. Measure end-of-turn latency separately from transcript latency. And report p50, p95, and p99 rather than single runs. It also warns against using response timestamps for millisecond timing. Capture wall-clock times on your own client instead.
Our guides on endpointing and streaming vs batch STT cover the tradeoffs in detail.
Hallucinated insertions
Some models invent words during silence, hold music, or noise. A caller breathes, and the transcript says "thank you". The agent then answers a question nobody asked.
Include silence-only and noise-only clips in every test set. Their reference is empty. Any output word is a hallucinated insertion. Count them separately, because WER cannot divide by an empty reference.
Worked example: low WER, wrong account number
Here is an illustrative 32-word caller turn. The STT output has three substituted words.

| Check | Result (illustrative) | Verdict |
|---|---|---|
| Raw WER | 3 errors / 32 words = 9.4% | Passes a 10% bar |
| WER after Whisper normalization | 2 errors / 25 tokens = 8.0% | Still passes |
| Account number "47192038" | Heard as "47192083" | Fail |
| Email "dana.kim@northwind.com" | Heard as "donna.kim@..." | Fail |
| Entity accuracy | 0 of 2 | Call cannot finish |
Notice the second row. The normalizer collapses eight digit words into one token. The error count falls, and WER looks even better. Yet the agent will look up the wrong account and email the wrong person.
Across a 400-utterance test set, this call barely moves corpus WER. Entity accuracy flags it immediately. That is why entity accuracy sits beside WER, not beneath it.
Normalization rules: the step that decides your WER
Normalization turns reference and hypothesis into comparable text. Without it, you score style differences as errors.
Take this pair:
| Version | Text |
|---|---|
| Reference | "I'd like to pay $250 on the 3rd." |
| STT output | "i would like to pay two hundred fifty dollars on the third" |
| Raw WER | 87.5% |
| Both after normalization | "i would like to pay $250 on the 3rd" |
| Normalized WER | 0% |
We computed this with jiwer and the Whisper English text normalizer. OpenAI's Whisper normalizer source lowercases text and drops bracketed words. It removes fillers like "um" and "uh", and expands contractions. It also converts spoken numbers and currency to digits, and maps British to American spelling.
Adopt these rules for your own STT evaluation:
- Apply the identical normalizer to references and every provider's output.
- Normalize for scoring only. Keep raw transcripts for entity and formatting checks.
- Decide on fillers deliberately. Removing "um" is fine for WER, but disfluencies can matter for turn-taking.
- Version your normalizer. A normalizer change can move WER more than a model change.
- Log the normalized strings. When a number looks off, you need to see what was compared.
Never let normalization hide entity errors. "Four seven one nine" and "4719" should match. "4719" and "4791" never should.
How do you test STT for accents, noise, and telephony audio?
This is where STT accuracy is won or lost. Models degrade sharply outside their training distribution, and real callers live outside it. Build the test set as a matrix of cohorts.
| Cohort dimension | Why it matters | How to source it |
|---|---|---|
| Accents and dialects | Accuracy gaps between speaker groups | Recruit speakers or sample consented production calls |
| Background noise | Street, car, office, TV, crosstalk | Record in place, or mix noise at set levels |
| Telephony audio | 8 kHz narrowband and lossy codecs | Route test audio through your real phone path |
| Domain vocabulary | Product names, jargon, IDs | Script turns from real call logs |
| Code-switching | Callers mixing languages mid-sentence | Native speakers reading realistic mixed turns |
| Silence and noise only | Hallucinated insertions | Recorded line noise, hold music, breathing |
Telephony deserves special care. Google Cloud's speech-to-text best practices note telephony is commonly 8,000 Hz. They advise sending that native rate rather than resampling. They also say lossy codecs and noise-reduction preprocessing can reduce accuracy. So test through the same codec and path your agent uses in production.
Report WER per cohort, never as one blend. For more depth, see our guides to accent handling, noise robustness, and testing STT in background noise. For global rollouts, the multilingual STT guide covers code-switching.
How many test utterances do you need?
Enough that a real difference is not noise. Entity accuracy behaves like a pass rate, so a simple binomial margin gives a useful floor.
The 95% margin of error is roughly 1.96 × √(p(1 − p) / n). At 95% entity accuracy:
| Entities per cohort | Approximate margin (illustrative) |
|---|---|
| 100 | ±4.3 points |
| 400 | ±2.1 points |
| 1,000 | ±1.4 points |
So a two-point difference between providers needs about 400 entities per cohort to mean much. WER is noisier, because errors cluster by speaker and utterance. Bootstrap by resampling whole utterances, or better, whole speakers, to get a confidence interval.
A practical starting point is 200 to 500 utterances per cohort. Use at least 20 speakers per cohort, so one voice cannot dominate. Keep a smaller smoke set of about 50 utterances for every commit.
How to run an STT evaluation step by step

1. Define the decisions. List what the transcript drives: intents, tool arguments, and verification steps. These define your critical entities.
2. Build the cohort matrix. Choose accents, noise levels, channels, and vocabulary that match your callers. Add silence-only clips.
3. Write verified references. Have humans transcribe each clip, with a second reviewer on entity values. Tag entities and keyterms.
4. Fix the normalizer. Pick one normalizer, version it, and apply it to every transcript.
5. Stream audio at real time. Send audio in real-time-sized chunks, as your agent does. Log every interim and final event with wall-clock timestamps.
6. Score per cohort. Compute WER, CER, entity accuracy, keyterm recall, stability, latency percentiles, and hallucinations.
7. Set gates. Choose pass bars from the cost of each error. Gate releases on the weakest cohort.
8. Re-run on every change. Model version, vendor setting, codec, or prompt hint changes all trigger the suite.
A Python STT evaluation harness
Below is a simplified harness. It is illustrative, not production code. It reads a JSONL file of results and prints metrics per cohort. Each row holds the reference, the final hypothesis, tagged entities, keyterms, and logged streaming events. We tested it with jiwer and a port of the Whisper normalizer.
pip install jiwer openai-whisper
# lighter alternative for the normalizer only:
pip install jiwer whisper-normalizer{"id": "c-0412", "cohort": "en-IN_8k_street",
"reference": "... my account number is four seven one nine two zero three eight ...",
"hypothesis": "... my account number is four seven one nine two zero eight three ...",
"entities": [{"kind": "account_number", "value": "47192038"},
{"kind": "email", "value": "dana.kim@northwind.com"}],
"keyterms": ["northwind"],
"speech_start": 0.30, "speech_end": 9.80,
"events": [{"t": 0.61, "is_final": false, "text": "hi i"},
{"t": 10.12, "is_final": true, "text": "..."}]}# stt_eval.py - simplified STT evaluation harness (illustrative, not production code)
import json
import re
import statistics
import sys
from collections import defaultdict
import jiwer
try:
from whisper.normalizers import EnglishTextNormalizer
except ImportError:
from whisper_normalizer.english import EnglishTextNormalizer
normalize = EnglishTextNormalizer()
def spoken_email(text):
# "dana dot kim at northwind dot com" -> "dana.kim@northwind.com"
t = text.lower().replace(" dot ", ".").replace(" at ", "@")
return re.sub(r"\s+", "", t)
def entity_hit(kind, expected, hyp_raw):
hyp_norm = normalize(hyp_raw)
if kind in ("account_number", "phone", "zip"):
return expected in re.sub(r"\D", "", hyp_norm)
if kind == "email":
return expected.lower() in spoken_email(hyp_raw)
return normalize(expected) in hyp_norm # names, product terms
def timing(events, speech_start, speech_end):
partials = [e for e in events if not e["is_final"]]
finals = [e for e in events if e["is_final"]]
ttfp = partials[0]["t"] - speech_start if partials else None
after_end = [e["t"] for e in finals if e["t"] >= speech_end]
ttf = after_end[0] - speech_end if after_end else None
# Stability: share of partial updates that rewrote an earlier word
rewrites, prev = 0, []
for e in partials:
words = e["text"].split()
if words[: len(prev)] != prev:
rewrites += 1
prev = words
churn = rewrites / len(partials) if partials else 0.0
return ttfp, ttf, churn
def p95(values):
values = sorted(values)
return values[int(0.95 * (len(values) - 1))] if values else None
def main(path):
rows = [json.loads(line) for line in open(path)]
by_cohort = defaultdict(list)
for r in rows:
by_cohort[r["cohort"]].append(r)
for cohort, items in sorted(by_cohort.items()):
speech = [r for r in items if r["reference"].strip()]
silence = [r for r in items if not r["reference"].strip()]
refs = [normalize(r["reference"]) for r in speech]
hyps = [normalize(r["hypothesis"]) for r in speech]
wer = jiwer.wer(refs, hyps) if speech else float("nan")
ent = [entity_hit(e["kind"], e["value"], r["hypothesis"])
for r in speech for e in r.get("entities", [])]
keys = [normalize(k) in normalize(r["hypothesis"])
for r in speech for k in r.get("keyterms", [])]
# Any words on a silence or noise-only clip are hallucinated insertions
halluc = sum(1 for r in silence if normalize(r["hypothesis"]).strip())
ttfps, ttfs, churns = [], [], []
for r in speech:
a, b, c = timing(r["events"], r["speech_start"], r["speech_end"])
if a is not None:
ttfps.append(a)
if b is not None:
ttfs.append(b)
churns.append(c)
print(f"[{cohort}] n={len(items)}")
print(f" WER {wer:.1%}")
if ent:
print(f" entity accuracy {sum(ent) / len(ent):.1%} ({sum(ent)}/{len(ent)})")
if keys:
print(f" keyterm recall {sum(keys) / len(keys):.1%}")
if silence:
print(f" hallucinations {halluc}/{len(silence)} silence clips")
if ttfps:
print(f" first partial p50 {statistics.median(ttfps) * 1000:.0f} ms")
if ttfs:
print(f" final after end p50 {statistics.median(ttfs) * 1000:.0f} ms"
f" / p95 {p95(ttfs) * 1000:.0f} ms")
print(f" partial churn {statistics.mean(churns):.1%}")
if __name__ == "__main__":
main(sys.argv[1])On the worked example plus one silence clip, it prints this:
[en-IN_8k_street] n=2
WER 8.0%
entity accuracy 0.0% (0/2)
keyterm recall 100.0%
hallucinations 1/1 silence clips
first partial p50 310 ms
final after end p50 320 ms / p95 320 ms
partial churn 33.3%A few caveats. The digit check only confirms the expected string appears somewhere in the turn. Real harnesses align entities by position. The stability metric is crude and ignores timing. And `speech_start` and `speech_end` should come from labeled audio or a VAD pass on the source file.
How do you benchmark STT providers fairly?
Hold the audio constant and vary only the provider. Then remove every other source of bias you can.
- Same audio, same path. Send identical files through identical codecs and sample rates.
- Same normalizer. Score every provider's output with one normalizer version.
- Each provider's best config. Enable its streaming mode, formatting, and vocabulary hints as you would in production.
- Real-time streaming. Measure latency from your region, with your chunk size, at realistic concurrency.
- Blind review. Hide provider names when humans judge semantic errors.
- Per-cohort results. A provider can win overall and lose your largest cohort.
Public leaderboards help you build a shortlist. The Open ASR Leaderboard standardizes WER and inverse real-time factor across many datasets. Its paper compares 86 open-source and proprietary systems across 12 datasets. Its evaluation code is public, so you can reuse the scoring approach. But its datasets are not your callers. Treat a leaderboard rank as a hypothesis, then test it on your traffic.
The same applies to single vendors. If you want to evaluate transcription accuracy in Deepgram, run this harness on your own calls. Our Deepgram STT testing guide and Deepgram latency and stability guide walk through that setup. For a wider shortlist, see our STT providers roundup.
Weighing cost against accuracy
Price per minute is rarely the real cost. Errors cost more. Estimate the cost of each failure type, then weigh it against price.
An illustrative model looks like this. A misheard account number forces a repeat or a human transfer. Suppose a transfer costs a few dollars and happens on 2% more calls with a cheaper model. That gap can outweigh a large per-minute saving at volume.
Price features separately too. Diarization, formatting, keyterms, and redaction can change the bill and the accuracy. Our STT cost vs accuracy guide walks through the tradeoff with worked numbers.
Regression testing when vendors update models
Vendors ship model updates, and your numbers can move overnight. Treat STT like any dependency with a changing version.
- Pin versions where possible. Deepgram, for example, documents a model version option. Check your vendor's docs for equivalents.
- Keep a frozen golden set. Never edit it casually. Add new cases to a separate set, then promote them deliberately.
- Diff, do not just score. Compare per-cohort metrics against the last accepted baseline. Flag any cohort that drops past a threshold.
- Replay production samples. Weekly, re-transcribe a consented sample of recent calls to catch ASR drift in live traffic.
- Watch proxies in production. Rising repeat requests, "sorry, can you say that again" turns, and transfer rates often signal STT drift first.
Our guides to voice agent metric drift and LLM update regressions cover the monitoring side.
How do you reduce word error rate in voice agent transcription?
Measure first, then fix the biggest cohort gap. Common levers, roughly in order of effort:
1. Send the native sample rate and a lossless or recommended codec.
2. Add keyterms or vocabulary hints for product names and jargon.
3. Tune endpointing so words are not cut off at turn ends.
4. Switch to a model built for telephony or conversational audio.
5. Constrain entity capture, for example asking callers to read digits in groups.
6. Confirm critical values back to the caller before acting on them.
Re-run the suite after each change, since fixes can trade one cohort for another. Our how to improve WER guide covers each lever in depth.
Common STT evaluation failure modes
| Failure mode | What it looks like | Fix |
|---|---|---|
| Clean-audio testing | Great lab WER, poor live results | Test through real telephony and noise |
| Blended scores | One accent cohort hidden at high WER | Report and gate per cohort |
| Mismatched normalization | Providers ranked by formatting style | One versioned normalizer for all |
| Finals-only scoring | Agent acts on wrong partials | Score stability and time to final |
| No silence clips | Invented "thank you" turns | Add silence and noise-only clips |
| Stale test set | Suite no longer matches callers | Refresh from production quarterly |
| One-time pass | Silent regression after vendor update | Re-run on every change and weekly |
What is a good STT accuracy for voice agents?
There is no universal number. It depends on the audio and on what the words drive.
Public benchmarks on read or prepared speech can show low single-digit WER for strong models. Real phone calls with accents, noise, and 8 kHz audio usually score worse. How much worse varies too much by domain to quote one figure. Measure your own baseline.
Set targets by consequence. A booking agent can tolerate more general transcription noise than one handling account numbers. For identifiers, aim for near-perfect entity accuracy and confirm values back to the caller. For general speech, aim for WER parity across cohorts, with no cohort far behind your best.
Where STT evaluation fits in full-call testing
A transcript can be slightly wrong in a way WER barely notices, yet break the call. The reverse also happens. A transcript with harmless errors still completes the task. Isolated STT scores cannot tell these apart.
That is why STT evaluation should connect to call outcomes. Our voice agent stack guide shows how the STT layer feeds everything downstream.
This is where Evalgent fits, as an independent third-party evaluation platform. Evalgent runs realistic calls and measures STT as one signal in context. Scenarios reproduce noisy, accented, and jargon-heavy calls. Profiles vary caller voices, so accuracy is reported per cohort. Metrics track WER, entity accuracy, and latency alongside task completion. Evaluations run these as automated batches of synthetic callers. Reviews let you hear the audio behind a bad number.
The transcript score tells you the words. The call result tells you whether they mattered. For the full discipline, see the AI voice agent testing pillar.
Frequently asked questions
What is STT evaluation?
STT evaluation is the practice of measuring a speech-to-text system's accuracy, latency, and robustness on audio that matches production. For voice agents it goes beyond word error rate. It also scores entity accuracy, keyterm recall, partial stability, time to final, endpointing, and hallucinated insertions. Every metric is reported per caller cohort, such as accent, noise level, and phone channel.
How do you evaluate transcription accuracy?
Compare each transcript to a human-verified reference after applying one shared normalizer. Compute word error rate as substitutions, deletions, and insertions divided by reference words. Then check tagged entities, like account numbers and emails, for exact matches. Libraries such as jiwer handle the alignment. Use audio that matches your callers, not a clean benchmark set.
What metrics are used for STT evaluation?
The core metrics are word error rate and character error rate. Voice agents also need entity accuracy, keyterm recall, formatting accuracy, and semantic error rate. Streaming adds partial stability, time to first partial, time to final, and endpointing accuracy. Finally, count hallucinated insertions on silence-only clips. Track them together, since one strong metric can hide a fatal weakness in another.
How do you reduce word error rate in voice agent transcription?
Measure per cohort first, then fix the largest gap. Send native sample rates and good codecs. Add keyterms or vocabulary hints for domain terms. Tune endpointing so words are not cut off. Try models built for telephony audio. For critical values, confirm them back to the caller. Re-run the full suite after each change, because fixes can shift errors between cohorts.
How do you benchmark STT providers fairly?
Hold the audio constant and vary only the provider. Use identical files, codecs, and sample rates, and one shared normalizer. Give each provider its best production configuration, including streaming and vocabulary hints. Measure latency from your own region at realistic concurrency. Report results per cohort. Public leaderboards like the Open ASR Leaderboard help shortlist, but your own traffic decides.
How many test utterances do you need for STT evaluation?
Start with 200 to 500 utterances per cohort and at least 20 speakers each. At 95% entity accuracy, 400 entities give roughly a two-point margin of error. That is enough to trust a two-point gap between providers. Bootstrap WER by resampling whole speakers for confidence intervals. Keep a smaller smoke set of about 50 utterances for every commit.
What is a good WER for voice agents?
There is no universal number. Public benchmarks on prepared speech can show low single-digit WER, but phone calls with accents and noise usually score worse. Set targets by consequence instead. Aim for near-perfect entity accuracy on identifiers, and confirm them with the caller. For general speech, aim for parity across cohorts, with no group far behind your best.
How do you monitor voice agents for ASR drift?
Keep a frozen golden test set and re-run it on every vendor model change. Pin model versions where your vendor allows it. Weekly, re-transcribe a consented sample of production calls and compare per-cohort metrics with your baseline. In production, watch proxies such as repeat requests, re-prompts, and transfer rates. Those often reveal STT drift before a scheduled test does.
The bottom line
STT evaluation for voice agents measures entity accuracy, latency, and per-cohort WER on production-like audio, not one blended score. The teams that ship reliable agents normalize consistently, test through real telephony, and re-run the suite on every model change.
Want to see your STT measured inside real calls, per cohort, with entity accuracy and latency? Book a demo.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: how to detect regressions from model updates you don't control
Voice agent regression testing for changes you don't control: pin each layer, run a nightly pinned-vs-latest canary, and use paired stats to detect drops.
Read more