Evalgent
Back to Blog
Voice AI Evaluation

STT evaluation: how to test speech-to-text for voice agents (2026)

Deepesh Jayal
17 min read
STT evaluation: how to test speech-to-text for voice agents (2026)
On this page

Speech-to-text is the first stage of every voice agent. It sets the ceiling for everything after it. If the transcript is wrong, the language model reasons over the wrong words. No prompt can fix that.

This guide is for engineers building voice agents on stacks like Pipecat, LiveKit, or Vapi. It covers the metrics, the test-set design, the normalization rules, and a working Python harness. It also covers fair provider comparisons and regression tests for vendor model updates. All example numbers are illustrative unless a source is linked.

8,000 Hz
Typical telephony sample rate, per Google Cloud guidance
86
ASR systems compared on the Open ASR Leaderboard paper
10 ms
Deepgram's default endpointing silence for streaming
2,400 ms
AssemblyAI's default max turn silence fallback

What is STT evaluation?

STT evaluation: the practice of measuring a speech-to-text system's accuracy, latency, and robustness on audio that matches production, broken down by caller cohort.

Speech to text evaluation answers two questions. How often does the system get the words right? And under what conditions does it fail? A model can look excellent on a clean benchmark and collapse on a noisy phone line. Your job is to find that gap before callers do.

The output is never one number. "The model is 95% accurate" predicts little. "It scores 95% on clean audio and 78% on accented calls in street noise" predicts production. That second sentence is what STT evaluation produces. If you are unsure about the terms, our STT vs ASR guide explains why they mean the same thing in practice.

Why generic WER misleads for voice agents

Word error rate is the standard accuracy metric. It is necessary. It is not sufficient for voice agents, for four reasons.

First, WER weights every word equally. "The" and the third digit of an account number count the same. A voice agent fails on the digit, not on "the".

Second, WER ignores time. A streaming agent acts on partial transcripts and endpoint signals. A model can score low WER on finals yet flip partials constantly, or finalize late.

Third, WER depends on normalization. The same transcript can score 87.5% or 0% depending on text rules. We show that exact case below.

Fourth, averages hide cohorts. A 7% blended WER can contain an accent group at 25%. Those callers churn while the dashboard stays green.

Entity accuracy: the share of critical values, such as account numbers, dates, names, emails, and addresses, that the transcript captures exactly. One wrong character counts as a miss.

How do you evaluate speech to text accuracy?

You compare the system's transcript to a human-verified reference. Then you count the edits needed to turn one into the other.

WER = (S + D + I) / N

S is substitutions, D is deletions, I is insertions, and N is the number of reference words. Lower is better. Because insertions count, WER can exceed 100%. The open-source jiwer library computes WER, CER, and alignments, and it shows each substitution, insertion, and deletion.

Character error rate applies the same formula to characters. It is more sensitive to small errors in names and IDs. Our WER vs CER guide covers when each one fits.

The math is the easy part. The hard part is the reference audio. It must represent your callers, with their accents, noise, lines, and vocabulary. Without that, you measure speech recognition on the wrong distribution.

The STT evaluation metrics that matter for voice agents

A complete assessment tracks accuracy, meaning, and time together. The table lists the core set. Pass bars are illustrative starting points. Set your own from the cost of each error.

MetricWhat it catchesHow to computeIllustrative pass bar
WEROverall transcription errors(S + D + I) / N on normalized textWithin 2 points of your best cohort
CERSmall errors in names and IDsSame formula at character levelTrack trend per cohort
Entity accuracyWrong digits, names, emails, datesExact match of tagged values98% or higher on IDs
Keyterm recallMissed product names and jargonShare of expected keyterms present95% or higher
Formatting accuracy"fifteen" vs "15", "$50" vs "fifty dollars"Compare formatted output to specMatches downstream parser
Semantic error rateErrors that change intentLabel or judge meaning-changing errorsUnder 2% of utterances
Partial stabilityInterim text rewriting itselfShare of partials that edit earlier wordsLow and stable across releases
Time to first partialSlow streaming startFirst interim time minus speech startp50 under 400 ms
Time to finalSlow turn completionFinal time minus speech endp95 within your turn budget
Endpointing accuracyCut-offs and late turn endsCompare detected end to labeled endFew premature cut-offs
Hallucinated insertionsWords invented on silence or noiseAny words on silence-only clipsNear zero

Accuracy and meaning metrics

Entity accuracy is usually the metric that decides whether a call succeeds. Tag every critical value in your references. Then check for an exact match after light normalization. Our STT entity accuracy guide goes deeper on tagging schemes.

Keyterm recall checks whether domain vocabulary survives. Product names, drug names, and plan tiers are common failures. Most vendors now accept vocabulary hints. Deepgram documents keyterm prompting. OpenAI's speech-to-text guide describes `prompt` and `keywords` parameters. Test recall with and without hints, because hints can also cause false insertions.

Formatting accuracy matters when a parser reads the transcript. If your tool call expects "15" and the STT returns "fifteen", the call fails. Test smart formatting settings as part of the model config.

Semantic error rate counts errors that change meaning. "Cancel" heard as "can sell" is one substitution. It is also a completely different intent. Label these by hand on a sample, or use a judge model with spot checks.

Streaming, latency, and endpointing metrics

Voice agents consume streams, not files. So measure the stream.

Most streaming APIs send interim results that later get corrected. Deepgram's interim results docs explain that early guesses near the audio edge are more likely wrong. Its responses mark finality with `is_final`. Its endpointing feature sets `speech_final` after a silence threshold. That threshold defaults to 10 ms and is configurable.

Endpointing now often mixes silence and semantics. AssemblyAI's turn detection docs describe a model-based trigger plus a silence fallback. Defaults include a 400 ms minimum silence when confident and a 2,400 ms maximum silence. Test the settings you will ship, not the defaults.

Deepgram's guide to measuring streaming latency makes two useful points. Measure end-of-turn latency separately from transcript latency. And report p50, p95, and p99 rather than single runs. It also warns against using response timestamps for millisecond timing. Capture wall-clock times on your own client instead.

Our guides on endpointing and streaming vs batch STT cover the tradeoffs in detail.

Hallucinated insertions

Some models invent words during silence, hold music, or noise. A caller breathes, and the transcript says "thank you". The agent then answers a question nobody asked.

Include silence-only and noise-only clips in every test set. Their reference is empty. Any output word is a hallucinated insertion. Count them separately, because WER cannot divide by an empty reference.

Worked example: low WER, wrong account number

Here is an illustrative 32-word caller turn. The STT output has three substituted words.

Why WER alone misleads: an illustrative transcript with a low word error rate that still gets the account number and email wrong, failing entity accuracy
CheckResult (illustrative)Verdict
Raw WER3 errors / 32 words = 9.4%Passes a 10% bar
WER after Whisper normalization2 errors / 25 tokens = 8.0%Still passes
Account number "47192038"Heard as "47192083"Fail
Email "dana.kim@northwind.com"Heard as "donna.kim@..."Fail
Entity accuracy0 of 2Call cannot finish

Notice the second row. The normalizer collapses eight digit words into one token. The error count falls, and WER looks even better. Yet the agent will look up the wrong account and email the wrong person.

Across a 400-utterance test set, this call barely moves corpus WER. Entity accuracy flags it immediately. That is why entity accuracy sits beside WER, not beneath it.

Normalization rules: the step that decides your WER

Normalization turns reference and hypothesis into comparable text. Without it, you score style differences as errors.

Take this pair:

VersionText
Reference"I'd like to pay $250 on the 3rd."
STT output"i would like to pay two hundred fifty dollars on the third"
Raw WER87.5%
Both after normalization"i would like to pay $250 on the 3rd"
Normalized WER0%

We computed this with jiwer and the Whisper English text normalizer. OpenAI's Whisper normalizer source lowercases text and drops bracketed words. It removes fillers like "um" and "uh", and expands contractions. It also converts spoken numbers and currency to digits, and maps British to American spelling.

Adopt these rules for your own STT evaluation:

  • Apply the identical normalizer to references and every provider's output.
  • Normalize for scoring only. Keep raw transcripts for entity and formatting checks.
  • Decide on fillers deliberately. Removing "um" is fine for WER, but disfluencies can matter for turn-taking.
  • Version your normalizer. A normalizer change can move WER more than a model change.
  • Log the normalized strings. When a number looks off, you need to see what was compared.

Never let normalization hide entity errors. "Four seven one nine" and "4719" should match. "4719" and "4791" never should.

How do you test STT for accents, noise, and telephony audio?

This is where STT accuracy is won or lost. Models degrade sharply outside their training distribution, and real callers live outside it. Build the test set as a matrix of cohorts.

Cohort dimensionWhy it mattersHow to source it
Accents and dialectsAccuracy gaps between speaker groupsRecruit speakers or sample consented production calls
Background noiseStreet, car, office, TV, crosstalkRecord in place, or mix noise at set levels
Telephony audio8 kHz narrowband and lossy codecsRoute test audio through your real phone path
Domain vocabularyProduct names, jargon, IDsScript turns from real call logs
Code-switchingCallers mixing languages mid-sentenceNative speakers reading realistic mixed turns
Silence and noise onlyHallucinated insertionsRecorded line noise, hold music, breathing

Telephony deserves special care. Google Cloud's speech-to-text best practices note telephony is commonly 8,000 Hz. They advise sending that native rate rather than resampling. They also say lossy codecs and noise-reduction preprocessing can reduce accuracy. So test through the same codec and path your agent uses in production.

Report WER per cohort, never as one blend. For more depth, see our guides to accent handling, noise robustness, and testing STT in background noise. For global rollouts, the multilingual STT guide covers code-switching.

How many test utterances do you need?

Enough that a real difference is not noise. Entity accuracy behaves like a pass rate, so a simple binomial margin gives a useful floor.

The 95% margin of error is roughly 1.96 × √(p(1 − p) / n). At 95% entity accuracy:

Entities per cohortApproximate margin (illustrative)
100±4.3 points
400±2.1 points
1,000±1.4 points

So a two-point difference between providers needs about 400 entities per cohort to mean much. WER is noisier, because errors cluster by speaker and utterance. Bootstrap by resampling whole utterances, or better, whole speakers, to get a confidence interval.

A practical starting point is 200 to 500 utterances per cohort. Use at least 20 speakers per cohort, so one voice cannot dominate. Keep a smaller smoke set of about 50 utterances for every commit.

How to run an STT evaluation step by step

An STT evaluation pipeline for voice agents: labeled test audio, transcription, text normalization, then scoring WER, entity accuracy, keyterm recall, stability and latency

1. Define the decisions. List what the transcript drives: intents, tool arguments, and verification steps. These define your critical entities.

2. Build the cohort matrix. Choose accents, noise levels, channels, and vocabulary that match your callers. Add silence-only clips.

3. Write verified references. Have humans transcribe each clip, with a second reviewer on entity values. Tag entities and keyterms.

4. Fix the normalizer. Pick one normalizer, version it, and apply it to every transcript.

5. Stream audio at real time. Send audio in real-time-sized chunks, as your agent does. Log every interim and final event with wall-clock timestamps.

6. Score per cohort. Compute WER, CER, entity accuracy, keyterm recall, stability, latency percentiles, and hallucinations.

7. Set gates. Choose pass bars from the cost of each error. Gate releases on the weakest cohort.

8. Re-run on every change. Model version, vendor setting, codec, or prompt hint changes all trigger the suite.

A Python STT evaluation harness

Below is a simplified harness. It is illustrative, not production code. It reads a JSONL file of results and prints metrics per cohort. Each row holds the reference, the final hypothesis, tagged entities, keyterms, and logged streaming events. We tested it with jiwer and a port of the Whisper normalizer.

pip install jiwer openai-whisper
# lighter alternative for the normalizer only:
pip install jiwer whisper-normalizer
{"id": "c-0412", "cohort": "en-IN_8k_street",
 "reference": "... my account number is four seven one nine two zero three eight ...",
 "hypothesis": "... my account number is four seven one nine two zero eight three ...",
 "entities": [{"kind": "account_number", "value": "47192038"},
              {"kind": "email", "value": "dana.kim@northwind.com"}],
 "keyterms": ["northwind"],
 "speech_start": 0.30, "speech_end": 9.80,
 "events": [{"t": 0.61, "is_final": false, "text": "hi i"},
            {"t": 10.12, "is_final": true, "text": "..."}]}
# stt_eval.py - simplified STT evaluation harness (illustrative, not production code)
import json
import re
import statistics
import sys
from collections import defaultdict

import jiwer

try:
    from whisper.normalizers import EnglishTextNormalizer
except ImportError:
    from whisper_normalizer.english import EnglishTextNormalizer

normalize = EnglishTextNormalizer()


def spoken_email(text):
    # "dana dot kim at northwind dot com" -> "dana.kim@northwind.com"
    t = text.lower().replace(" dot ", ".").replace(" at ", "@")
    return re.sub(r"\s+", "", t)


def entity_hit(kind, expected, hyp_raw):
    hyp_norm = normalize(hyp_raw)
    if kind in ("account_number", "phone", "zip"):
        return expected in re.sub(r"\D", "", hyp_norm)
    if kind == "email":
        return expected.lower() in spoken_email(hyp_raw)
    return normalize(expected) in hyp_norm  # names, product terms


def timing(events, speech_start, speech_end):
    partials = [e for e in events if not e["is_final"]]
    finals = [e for e in events if e["is_final"]]
    ttfp = partials[0]["t"] - speech_start if partials else None
    after_end = [e["t"] for e in finals if e["t"] >= speech_end]
    ttf = after_end[0] - speech_end if after_end else None
    # Stability: share of partial updates that rewrote an earlier word
    rewrites, prev = 0, []
    for e in partials:
        words = e["text"].split()
        if words[: len(prev)] != prev:
            rewrites += 1
        prev = words
    churn = rewrites / len(partials) if partials else 0.0
    return ttfp, ttf, churn


def p95(values):
    values = sorted(values)
    return values[int(0.95 * (len(values) - 1))] if values else None


def main(path):
    rows = [json.loads(line) for line in open(path)]
    by_cohort = defaultdict(list)
    for r in rows:
        by_cohort[r["cohort"]].append(r)

    for cohort, items in sorted(by_cohort.items()):
        speech = [r for r in items if r["reference"].strip()]
        silence = [r for r in items if not r["reference"].strip()]
        refs = [normalize(r["reference"]) for r in speech]
        hyps = [normalize(r["hypothesis"]) for r in speech]
        wer = jiwer.wer(refs, hyps) if speech else float("nan")

        ent = [entity_hit(e["kind"], e["value"], r["hypothesis"])
               for r in speech for e in r.get("entities", [])]
        keys = [normalize(k) in normalize(r["hypothesis"])
                for r in speech for k in r.get("keyterms", [])]
        # Any words on a silence or noise-only clip are hallucinated insertions
        halluc = sum(1 for r in silence if normalize(r["hypothesis"]).strip())

        ttfps, ttfs, churns = [], [], []
        for r in speech:
            a, b, c = timing(r["events"], r["speech_start"], r["speech_end"])
            if a is not None:
                ttfps.append(a)
            if b is not None:
                ttfs.append(b)
            churns.append(c)

        print(f"[{cohort}] n={len(items)}")
        print(f"  WER             {wer:.1%}")
        if ent:
            print(f"  entity accuracy {sum(ent) / len(ent):.1%} ({sum(ent)}/{len(ent)})")
        if keys:
            print(f"  keyterm recall  {sum(keys) / len(keys):.1%}")
        if silence:
            print(f"  hallucinations  {halluc}/{len(silence)} silence clips")
        if ttfps:
            print(f"  first partial   p50 {statistics.median(ttfps) * 1000:.0f} ms")
        if ttfs:
            print(f"  final after end p50 {statistics.median(ttfs) * 1000:.0f} ms"
                  f" / p95 {p95(ttfs) * 1000:.0f} ms")
        print(f"  partial churn   {statistics.mean(churns):.1%}")


if __name__ == "__main__":
    main(sys.argv[1])

On the worked example plus one silence clip, it prints this:

[en-IN_8k_street] n=2
  WER             8.0%
  entity accuracy 0.0% (0/2)
  keyterm recall  100.0%
  hallucinations  1/1 silence clips
  first partial   p50 310 ms
  final after end p50 320 ms / p95 320 ms
  partial churn   33.3%

A few caveats. The digit check only confirms the expected string appears somewhere in the turn. Real harnesses align entities by position. The stability metric is crude and ignores timing. And `speech_start` and `speech_end` should come from labeled audio or a VAD pass on the source file.

How do you benchmark STT providers fairly?

Hold the audio constant and vary only the provider. Then remove every other source of bias you can.

  • Same audio, same path. Send identical files through identical codecs and sample rates.
  • Same normalizer. Score every provider's output with one normalizer version.
  • Each provider's best config. Enable its streaming mode, formatting, and vocabulary hints as you would in production.
  • Real-time streaming. Measure latency from your region, with your chunk size, at realistic concurrency.
  • Blind review. Hide provider names when humans judge semantic errors.
  • Per-cohort results. A provider can win overall and lose your largest cohort.

Public leaderboards help you build a shortlist. The Open ASR Leaderboard standardizes WER and inverse real-time factor across many datasets. Its paper compares 86 open-source and proprietary systems across 12 datasets. Its evaluation code is public, so you can reuse the scoring approach. But its datasets are not your callers. Treat a leaderboard rank as a hypothesis, then test it on your traffic.

The same applies to single vendors. If you want to evaluate transcription accuracy in Deepgram, run this harness on your own calls. Our Deepgram STT testing guide and Deepgram latency and stability guide walk through that setup. For a wider shortlist, see our STT providers roundup.

Weighing cost against accuracy

Price per minute is rarely the real cost. Errors cost more. Estimate the cost of each failure type, then weigh it against price.

An illustrative model looks like this. A misheard account number forces a repeat or a human transfer. Suppose a transfer costs a few dollars and happens on 2% more calls with a cheaper model. That gap can outweigh a large per-minute saving at volume.

Price features separately too. Diarization, formatting, keyterms, and redaction can change the bill and the accuracy. Our STT cost vs accuracy guide walks through the tradeoff with worked numbers.

Regression testing when vendors update models

Vendors ship model updates, and your numbers can move overnight. Treat STT like any dependency with a changing version.

  • Pin versions where possible. Deepgram, for example, documents a model version option. Check your vendor's docs for equivalents.
  • Keep a frozen golden set. Never edit it casually. Add new cases to a separate set, then promote them deliberately.
  • Diff, do not just score. Compare per-cohort metrics against the last accepted baseline. Flag any cohort that drops past a threshold.
  • Replay production samples. Weekly, re-transcribe a consented sample of recent calls to catch ASR drift in live traffic.
  • Watch proxies in production. Rising repeat requests, "sorry, can you say that again" turns, and transfer rates often signal STT drift first.

Our guides to voice agent metric drift and LLM update regressions cover the monitoring side.

How do you reduce word error rate in voice agent transcription?

Measure first, then fix the biggest cohort gap. Common levers, roughly in order of effort:

1. Send the native sample rate and a lossless or recommended codec.

2. Add keyterms or vocabulary hints for product names and jargon.

3. Tune endpointing so words are not cut off at turn ends.

4. Switch to a model built for telephony or conversational audio.

5. Constrain entity capture, for example asking callers to read digits in groups.

6. Confirm critical values back to the caller before acting on them.

Re-run the suite after each change, since fixes can trade one cohort for another. Our how to improve WER guide covers each lever in depth.

Common STT evaluation failure modes

Failure modeWhat it looks likeFix
Clean-audio testingGreat lab WER, poor live resultsTest through real telephony and noise
Blended scoresOne accent cohort hidden at high WERReport and gate per cohort
Mismatched normalizationProviders ranked by formatting styleOne versioned normalizer for all
Finals-only scoringAgent acts on wrong partialsScore stability and time to final
No silence clipsInvented "thank you" turnsAdd silence and noise-only clips
Stale test setSuite no longer matches callersRefresh from production quarterly
One-time passSilent regression after vendor updateRe-run on every change and weekly

What is a good STT accuracy for voice agents?

There is no universal number. It depends on the audio and on what the words drive.

Public benchmarks on read or prepared speech can show low single-digit WER for strong models. Real phone calls with accents, noise, and 8 kHz audio usually score worse. How much worse varies too much by domain to quote one figure. Measure your own baseline.

Set targets by consequence. A booking agent can tolerate more general transcription noise than one handling account numbers. For identifiers, aim for near-perfect entity accuracy and confirm values back to the caller. For general speech, aim for WER parity across cohorts, with no cohort far behind your best.

Where STT evaluation fits in full-call testing

A transcript can be slightly wrong in a way WER barely notices, yet break the call. The reverse also happens. A transcript with harmless errors still completes the task. Isolated STT scores cannot tell these apart.

That is why STT evaluation should connect to call outcomes. Our voice agent stack guide shows how the STT layer feeds everything downstream.

This is where Evalgent fits, as an independent third-party evaluation platform. Evalgent runs realistic calls and measures STT as one signal in context. Scenarios reproduce noisy, accented, and jargon-heavy calls. Profiles vary caller voices, so accuracy is reported per cohort. Metrics track WER, entity accuracy, and latency alongside task completion. Evaluations run these as automated batches of synthetic callers. Reviews let you hear the audio behind a bad number.

The transcript score tells you the words. The call result tells you whether they mattered. For the full discipline, see the AI voice agent testing pillar.

Frequently asked questions

What is STT evaluation?

STT evaluation is the practice of measuring a speech-to-text system's accuracy, latency, and robustness on audio that matches production. For voice agents it goes beyond word error rate. It also scores entity accuracy, keyterm recall, partial stability, time to final, endpointing, and hallucinated insertions. Every metric is reported per caller cohort, such as accent, noise level, and phone channel.

How do you evaluate transcription accuracy?

Compare each transcript to a human-verified reference after applying one shared normalizer. Compute word error rate as substitutions, deletions, and insertions divided by reference words. Then check tagged entities, like account numbers and emails, for exact matches. Libraries such as jiwer handle the alignment. Use audio that matches your callers, not a clean benchmark set.

What metrics are used for STT evaluation?

The core metrics are word error rate and character error rate. Voice agents also need entity accuracy, keyterm recall, formatting accuracy, and semantic error rate. Streaming adds partial stability, time to first partial, time to final, and endpointing accuracy. Finally, count hallucinated insertions on silence-only clips. Track them together, since one strong metric can hide a fatal weakness in another.

How do you reduce word error rate in voice agent transcription?

Measure per cohort first, then fix the largest gap. Send native sample rates and good codecs. Add keyterms or vocabulary hints for domain terms. Tune endpointing so words are not cut off. Try models built for telephony audio. For critical values, confirm them back to the caller. Re-run the full suite after each change, because fixes can shift errors between cohorts.

How do you benchmark STT providers fairly?

Hold the audio constant and vary only the provider. Use identical files, codecs, and sample rates, and one shared normalizer. Give each provider its best production configuration, including streaming and vocabulary hints. Measure latency from your own region at realistic concurrency. Report results per cohort. Public leaderboards like the Open ASR Leaderboard help shortlist, but your own traffic decides.

How many test utterances do you need for STT evaluation?

Start with 200 to 500 utterances per cohort and at least 20 speakers each. At 95% entity accuracy, 400 entities give roughly a two-point margin of error. That is enough to trust a two-point gap between providers. Bootstrap WER by resampling whole speakers for confidence intervals. Keep a smaller smoke set of about 50 utterances for every commit.

What is a good WER for voice agents?

There is no universal number. Public benchmarks on prepared speech can show low single-digit WER, but phone calls with accents and noise usually score worse. Set targets by consequence instead. Aim for near-perfect entity accuracy on identifiers, and confirm them with the caller. For general speech, aim for parity across cohorts, with no group far behind your best.

How do you monitor voice agents for ASR drift?

Keep a frozen golden test set and re-run it on every vendor model change. Pin model versions where your vendor allows it. Weekly, re-transcribe a consented sample of production calls and compare per-cohort metrics with your baseline. In production, watch proxies such as repeat requests, re-prompts, and transfer rates. Those often reveal STT drift before a scheduled test does.

The bottom line

STT evaluation for voice agents measures entity accuracy, latency, and per-cohort WER on production-like audio, not one blended score. The teams that ship reliable agents normalize consistently, test through real telephony, and re-run the suite on every model change.

Want to see your STT measured inside real calls, per cohort, with entity accuracy and latency? Book a demo.

Related Articles