Evalgent
Back to Blog
Voice AI Evaluation

TTS evaluation for voice agents: beyond MOS to a phone-ready scorecard

Deepesh Jayal
Updated
25 min read
TTS evaluation for voice agents: beyond MOS to a phone-ready scorecard
On this page

Most TTS evaluation advice stops at "collect a MOS score and listen to some samples." That advice was built for research papers comparing synthesis systems on studio headphones. A voice agent has a different job. It reads text an LLM wrote a moment ago, full of phone numbers, dosages and confirmation codes, streams it in pieces, and pushes it through an 8 kHz telephone codec to someone standing in a parking lot.

This guide is for teams building or buying production voice agents. It covers what research says about MOS reliability, how to size listening tests, round-trip intelligibility, text normalization failures, telephony degradation, streaming latency and production tracing. It ends with a scorecard and test matrix you can copy.

1.28
MOS points one system lost when re-rated among only top systems (Cooper & Yamagishi, 2023)
0.31
Rank correlation of the top 11 systems across two listening-test contexts (same study)
3.21 vs 3.53
MOS for identical audio when asked to rate "naturalness" vs "quality" (Kirkland et al., 2023)
0.79
Best zero-shot system-level rank correlation for an automatic MOS predictor on French TTS (VoiceMOS 2023)

What TTS evaluation means for a voice agent

TTS evaluation: measuring whether a text-to-speech system delivers the right words, with usable timing and an acceptable voice, under the conditions your agent actually runs in.

For a voice agent that breaks into five dimensions. Each one fails independently, and each needs its own test.

DimensionQuestion it answersPrimary measurementWhat a demo clip hides
IntelligibilityCan a listener recover the words?Round-trip CER/WER through STT, plus listening checksStudio audio, short sentences
Entity fidelityWere numbers, dates, names and codes spoken correctly?Exact-match entity rateDemo scripts contain no live data
Latency and streamingHow soon does the caller hear audio, and does it flow?Caller-heard first audio, pause anomalies at chunk joinsPre-rendered files have no streaming
Telephony robustnessDoes it survive 8 kHz μ-law and transcoding?All of the above, re-run after the phone chainSamples are 24 or 44.1 kHz
Naturalness and consistencyDoes it sound acceptable and stay the same voice all call?Paired preference tests, speaker similarity per turnClips are seconds long

Naturalness sits last on purpose. Vendors are closest together on it, it is the hardest to measure, and its failures cost the least. A slightly flat voice annoys callers. A voice that reads "$1,250.50" as "one thousand two hundred fifty point five zero dollars" can trigger a dispute. Our mean opinion score guide covers the basics of the scale; the next section explains why you should not pick a voice with it.

Why MOS is unreliable for picking a voice

A mean opinion score asks listeners to rate clips on a 1-to-5 scale and averages the result. The method comes from ITU-T P.800, which was written to rate telephone transmission quality. Three findings from recent research explain why a MOS number from a vendor page, or even from your own test, should not decide a voice.

MOS is relative, not absolute

Cooper and Yamagishi (Interspeech 2023) took the BVCC dataset, which holds ratings for 187 TTS and voice conversion systems, and re-ran listening tests that "zoomed in" on progressively smaller groups of the best systems. Listeners kept using the whole 1-to-5 scale no matter how good the samples were. This is called range-equalizing bias.

The effects were large. One system dropped by 1.28 MOS points when it was rated only alongside top systems. The system-level Spearman correlation between the original ranking of the top 11 systems and their ranking in the narrowest test fell to 0.31. In plain terms, the ordering of good systems was close to reshuffled by changing which other systems were in the test. The same paper notes that reported MOS for Tacotron 2 went from 4.53 in its 2018 paper to 4.20 in a 2020 study and 3.01 in a 2022 study, as better systems entered the comparison pool.

Le Maguer, King and Harte (Computer Speech and Language, 2024) reached the same conclusion from a different angle. Replicating the 2013 Blizzard Challenge, they showed MOS behaves as a relative score and is sensitive to whether low and high anchors are present.

What this means in production: a MOS of 4.4 on one vendor's page and 4.1 on another's tell you nothing about which is better. Even inside your own test, adding a weak baseline voice changes how far apart your finalists look.

The instructions change the number

Kirkland et al. (SSW 2023) surveyed 133 TTS papers from Interspeech and SSW. Only 2 papers published the actual question asked of listeners, and more than 75 percent did not report the scale increments. They then ran the same stimuli through four test variants with 104 listeners. Asking listeners to rate "overall quality" produced an average MOS of 3.53. Asking them to rate "how natural it sounds" produced 3.21 for identical audio. For one system the gap was 2.26 versus 2.88.

The rankings held in that experiment, which is the useful takeaway. MOS can order systems that differ a lot. It is weak at separating close competitors, and modern commercial voices are close competitors.

Two-panel chart showing MOS context effects: rank correlation of top TTS systems falling from 0.83 to 0.31 as tests zoom in, and identical audio scoring differently under quality versus naturalness instructions

Automatic MOS predictors and their correlation limits

Automatic predictors such as UTMOS, DNSMOS and NISQA promise MOS without listeners. Each was trained for a specific target, and that matters more than the headline correlation.

  • UTMOS won several metrics in the VoiceMOS Challenge 2022, trained on BVCC. Cooper and Yamagishi note that BVCC audio is all 16 kHz.
  • DNSMOS was built, in its authors' words, to evaluate noise suppressors. It predicts how clean a recording sounds, not whether synthetic prosody is natural.
  • NISQA targets distortions in communication networks. It is closer to telephony but was not designed to judge synthesis.

The VoiceMOS Challenge 2023 tested predictors on domains they had never seen. In the 2022 out-of-domain track, with only 136 labeled samples from the target domain, the top system reached a system-level rank correlation of 0.979. In 2023's zero-shot French TTS track, the best team reached 0.79 and the average team 0.57. The UTMOS and SSL-MOS baselines, used unchanged, scored above 0.5 on one French sub-track but below 0.35 on the other, which had reverberant training audio. The organizers concluded that general-purpose MOS prediction is still an open problem.

Your production audio is further out of domain than any of those tracks: 8 kHz μ-law, LLM-written conversational text, and voices that did not exist when the predictor was trained. Use predictors for one job only. Run the same voice on the same scripts before and after a change, and treat a drop as a smoke alarm that triggers listening. Never use a predictor score to choose between two vendors.

Better listening tests: CMOS, MUSHRA and AB preference

If a single-stimulus MOS cannot separate close voices, compare them directly. Three designs are worth knowing.

DesignWhat listeners doStrengthWeaknessUse it for
ACR MOS (ITU-T P.800)Rate one clip at a time, 1 to 5Familiar, cheap per ratingContext effects, poor at close callsSanity checks only
CMOS / CCR (P.800 comparison)Hear A then B, rate B relative to A from -3 to +3Sensitive to small differencesOrder effects unless counterbalancedVersion-to-version regressions
MUSHRA (ITU-R BS.1534)Rate several systems side by side, 0 to 100, with hidden reference and anchorMany systems per trialAnchors itself bias scoresShortlisting 3 to 6 voices
AB forced-choice preferencePick the better of two clips of the same textSimplest to analyze, robustOne bit per judgmentFinal vendor decision

MUSHRA has its own trap for TTS. Rethinking MUSHRA (2024), which used 492 listeners rating Hindi and Tamil TTS, found a reference-matching bias: listeners penalize synthetic speech for differing from the human reference even when the synthetic sample is better. If your human reference is a single recorded agent, MUSHRA will reward voices that sound like that person, not voices callers prefer.

For a voice agent the most useful design is AB preference on your own scripts, played through the telephone chain, with the order of A and B randomized per trial. It answers the actual business question, "which of these two would callers rather hear," and it avoids most scale-use bias because there is no scale.

How many raters and clips you need

Listener count is where most in-house tests fail. Cooper and Yamagishi cite earlier work by Wester and colleagues on Blizzard 2013 data in which statistical significance only stabilized at about 30 listeners. Below that, rankings move around with every new rater.

For an AB preference test, the number of judgments needed to detect a true preference rate p1 against the null of 50 percent, at two-sided alpha 0.05 and 80 percent power, is:

n = [(1.96 × √(0.5 × 0.5) + 0.84 × √(p1 × (1 − p1))) / (p1 − 0.5)]²

  • To detect a 60/40 preference: n = [(0.98 + 0.84 × 0.490) / 0.10]² ≈ 194 judgments.
  • To detect a 55/45 preference: n = [(0.98 + 0.84 × 0.497) / 0.05]² ≈ 783 judgments.

Those counts assume independent judgments. Twenty judgments from one rater are not independent, because that rater has a taste. Inflate by the design effect, DEFF = 1 + (m − 1) × ICC, where m is judgments per rater and ICC is the within-rater correlation. As an illustrative assumption, with m = 20 and ICC = 0.05, DEFF = 1.95. The 60/40 test then needs about 380 judgments, or roughly 19 raters doing 20 pairs each. In practice, plan 25 to 30 raters to clear the 30-listener stability point after attention-check exclusions.

For a MOS comparison between two voices, the per-voice sample size is n = 2 × (1.96 + 0.84)² × σ² / Δ². Cooper and Yamagishi report per-rating standard deviations between about 0.8 and 1.2. With σ = 1.0, detecting a 0.2-point difference needs about 392 ratings per voice, and detecting 0.1 needs about 1,570. That is why MOS differences of a tenth of a point on vendor pages carry so little weight.

Practical rules from the papers: publish the exact question and scale labels; add attention checks, as Kirkland et al. did with spoken instructions inside clips; keep only the finalists in the pool; and rate audio after the phone chain, not on studio headphones.

The same sizing logic applies to call-level experiments; see A/B testing voice agents.

Round-trip intelligibility: TTS to STT as an objective proxy

Listening tests are slow, so every build also needs an automated check. Round-trip intelligibility synthesizes the text, transcribes the audio with an STT model, and scores the transcript against the input.

Research benchmarks use the same idea. The Seed-TTS evaluation set scores English intelligibility as WER using Whisper-large-v3 and scores speaker similarity as cosine similarity of WavLM speaker-verification embeddings. For a voice agent, adapt it in four ways.

1. Score after the phone chain. Render at 8 kHz μ-law before transcription, because that is what the caller receives.

2. Prefer CER and entity exact-match over WER. WER treats "fifteen" versus "fifty" as one word error out of twenty. For a dosage, that one word is the whole sentence. Our guide to word error rate covers the formula; for entities, use the approach in STT entity accuracy.

3. Use two STT engines from different vendors. A round-trip error can be the STT's fault. If both engines fail on the same span, the TTS is the likely cause. If only one does, listen before blaming the voice.

4. Normalize both sides the same way. Lowercase, strip punctuation and map number formats before scoring, or formatting differences will swamp real errors.

There is a subtle limit. Modern STT models carry strong language models that "repair" what they hear. If the TTS slurs "metoprolol" but the STT knows the drug name, the transcript can come back correct. Round-trip checks catch skipped words, wrong numbers and truncation, but under-count mispronunciations of familiar words, so pair them with human listening on domain terms.

Here is a simplified round-trip harness using Deepgram's REST APIs. The parameters come from Deepgram's TTS media output settings: `mulaw` supports `container=none` and 8000 Hz. The STT request declares the raw encoding and sample rate, which Deepgram requires for headerless audio.

# Simplified round-trip intelligibility check (illustrative).
# pip install requests jiwer
import os, re, requests, jiwer

DG = "https://api.deepgram.com/v1"
HEAD = {"Authorization": f"Token {os.environ['DEEPGRAM_API_KEY']}"}

def tts_phone(text, voice="aura-2-thalia-en"):
    # Raw 8 kHz mu-law, no WAV header: what a Twilio leg would carry.
    r = requests.post(f"{DG}/speak",
                      params={"model": voice, "encoding": "mulaw",
                              "sample_rate": 8000, "container": "none"},
                      headers={**HEAD, "Content-Type": "application/json"},
                      json={"text": text}, timeout=30)
    r.raise_for_status()
    return r.content

def stt(audio):
    r = requests.post(f"{DG}/listen",
                      params={"model": "nova-3", "encoding": "mulaw",
                              "sample_rate": 8000, "smart_format": "true"},
                      headers={**HEAD, "Content-Type": "audio/mulaw"},
                      data=audio, timeout=60)
    r.raise_for_status()
    return r.json()["results"]["channels"][0]["alternatives"][0]["transcript"]

def norm(s):
    return re.sub(r"[^a-z0-9 ]", "", s.lower()).strip()

def digits(s):
    return re.sub(r"\D", "", s)

case = {"text": "Your refill of 25 mg metoprolol ships Friday. "
                "Call 415-555-0134 with code 7Q4K.",
        "entities": {"phone": "4155550134", "dose": "25"}}

hyp = stt(tts_phone(case["text"]))
print("CER:", round(jiwer.cer(norm(case["text"]), norm(hyp)), 3))
for name, want in case["entities"].items():
    print(name, "PASS" if want in digits(hyp) else "FAIL", "|", hyp)

Two gotchas from this setup. First, `smart_format` turns spoken digits back into "415-555-0134," which is what you want for entity matching but means your WER reference must use the same formatting. Second, if you forget `container=none` when piping TTS audio into a telephony leg, Deepgram's docs warn that the WAV header gets played as audio and causes clicks. Twilio's Media Streams docs carry the same warning from the other side: the outbound payload must not contain file header bytes.

Entity rendering: where natural voices say the wrong thing

Entity errors are the TTS failures that cost money. They come from text normalization, the step that converts "Dr. Lee, 5/6, $40.10" into words before synthesis.

Sproat and Jaitly (2016) studied neural text normalization on a large corpus and found that models with high overall accuracy still produced errors that "would convey completely the wrong message," such as reading the wrong number or substituting one unit for another. Classic rule-based normalizers rarely make that kind of error. Their fix was a rule-based filter over the neural output. The lesson carries over to today's end-to-end TTS: overall accuracy can look excellent while rare, high-impact entity errors slip through.

Vendor documentation confirms where this bites. ElevenLabs' model documentation states that normalization is disabled by default for Flash v2.5 to keep latency low, and warns that phone numbers, dates and currencies may be read unclearly as a result. The recommended fixes are to have the LLM normalize text first or to set `apply_text_normalization` to "on," which is limited to Enterprise plans for v2.5 models. Deepgram's Aura-2 formatting guide recommends spelling letter codes with spaces ("J O H N") rather than hyphens or commas, and putting acronyms in quotes.

The practical rule: normalize in deterministic, unit-tested code between the LLM and the TTS, and treat the TTS normalizer as a second line of defense.

Failure taxonomy grid of eight entity types such as currency, dates, phone numbers and drug doses, showing written input, a typical wrong spoken rendering, and the correct rendering
Entity classWritten formTypical failure pattern (illustrative)Pass rule
Currency$1,250.50"one thousand two hundred fifty point five zero"Dollars and cents spoken, amount exact
Dates05/06/2026Read as June 5 for a US callerLocale-correct month and day
Phone numbers415-555-0134"four hundred fifteen, five hundred fifty-five..."Digits in groups of 3-3-4
Alphanumeric codes7Q4K"seven queue forty" or letters mergedEach character, with pauses
AbbreviationsDr., St., mg"Drive" for Doctor, "Street" for SaintContext-correct expansion
Emails and URLsj.smith@acme.co"jsmith acme co" with symbols dropped"dot" and "at" spoken
Drug and product namesmetoprololStress on wrong syllable, letters droppedMatches approved pronunciation
Units and ranges10-20 mg"ten twenty M G""ten to twenty milligrams"

Steps to validate pronunciation of domain terms like drug names

1. Build the term list from real data. Pull every drug, product, place and surname your agent can say from formularies, catalogs and CRM data, weighted by frequency.

2. Write the approved pronunciation for each term. For drugs, use a pharmacist or clinical reviewer, not a generic dictionary.

3. Render each term in three carrier sentences (start, middle and end of a sentence), because coarticulation changes pronunciation.

4. Run round-trip with two STT engines and flag any term either engine misses.

5. Have a domain reviewer listen to every term, including the ones that passed round-trip, since STT language models can mask errors on familiar words.

6. Fix with the vendor's lexicon tools. ElevenLabs supports pronunciation dictionaries on its WebSocket, and its real-time guide notes that phoneme-based dictionaries need `enable_ssml_parsing=true` on the connection URI.

7. Re-test after every model or voice change. Lexicon behavior can change between model versions.

EmergentTTS-Eval offers 1,645 hard test cases, including complex pronunciation (URLs, formulas) and foreign words, scored by a large audio language model. It is a good source of difficult sentences, but your own entity list is more relevant.

Telephony degradation: why a 24 kHz winner can lose on the phone

Vendor samples play at 24 kHz or higher. Most calls do not. Twilio's Media Streams documentation states that the stream encoding is always `audio/x-mulaw` at 8000 Hz, mono. Deepgram's streaming TTS defaults to linear16 at 24000 Hz. Somewhere between those two, your stack resamples.

Here is the mechanism, and why it changes rankings:

  • Bandwidth. At 8 kHz sampling the Nyquist limit is 4 kHz, and narrowband telephony is conventionally treated as about 300 to 3,400 Hz. Much of the energy that distinguishes "s" from "f" and "th" sits above that range. Voices whose clarity comes from crisp high-frequency detail lose their advantage. Spelled codes suffer most: "S as in Sam" exists because of this band limit.
  • Companding. G.711 μ-law stores each sample in 8 bits on a logarithmic curve. Breathy, low-level detail that sounds intimate at 24 kHz turns into quantization noise.
  • Resampling. Converting 24 kHz to 8 kHz requires a low-pass filter first. A naive decimator aliases high-frequency energy back into the band as a metallic edge. The resampler is chosen by your framework or transport, so two stacks using the same voice can sound different.
  • Transcoding chains. A WebRTC leg may carry Opus, then a gateway converts to G.711 for the PSTN, then the carrier may transcode again. Each hop removes something. See SIP vs WebRTC for where those hops occur.
  • Predictor blind spots. BVCC, the data behind several MOS predictors, is 16 kHz, so they never learned what good 8 kHz synthetic speech sounds like.

The protocol: render every candidate through the same phone chain you will deploy and run every evaluation in this guide on the degraded audio. Add a listener-side noise condition too. Callers in cars and kitchens hear your agent over their own noise, so mix the decoded output with babble at about 5 to 10 dB SNR and re-run round-trip CER.

Streaming latency: TTFB, chunk boundaries and prosody breaks

A voice agent never renders a whole file. It streams LLM text into the TTS and streams audio out. That creates two failure modes a file-based test cannot see: delay before the first audio, and audible seams where chunks join.

Three different TTFBs

"Time to first byte" means at least three different things, and dashboards rarely say which.

LevelStarts whenEnds whenWhere you read it
Model TTFBRequest reaches vendorFirst audio leaves modelVendor docs; ElevenLabs marks its figures "excluding application and network latency"
Service TTFBYour process sends textFirst audio bytes arrivePipecat `metrics.ttfb` on TTS spans; LiveKit `TTSMetrics.ttfb`
Caller-heard first audioText is ready in your pipelineAudio plays on the caller's lineTwilio `mark` events (returned when buffered audio finishes playing), or the gap in a stereo call recording

The ElevenLabs model page lists Flash v2.5 at about 75 ms and Eleven v4 Turbo at a median inference latency of about 100 ms, both footnoted as excluding application and network latency. Those are model numbers. The number callers feel is the third row, and the gap between rows one and three is where most "the TTS is slow" complaints actually live. Our time to first audio guide covers the full-turn view.

How buffering policy adds hidden delay

TTS models need some text context to produce good prosody, so streaming APIs buffer input. Each vendor exposes this differently, and the defaults interact with how your orchestrator sends text.

  • ElevenLabs uses `chunk_length_schedule`, default `[120, 160, 250, 290]`. Its real-time guide says audio is generated after 120 characters have been sent, then at the next thresholds. Sending `flush: true` forces generation of buffered text, and ElevenLabs advises flushing at the end of each conversational turn. Its `auto_mode` disables the schedule and buffers, but the API reference recommends it only for full sentences and warns that partial sentences degrade quality sharply.
  • Cartesia uses contexts. You send chunks with the same `context_id` and `continue: true` to keep prosody across inputs, and the contexts docs recommend `max_buffer_delay_ms` when streaming tokens. Contexts expire 1 second after the last audio output, so a slow tool call between chunks can break the context.
  • Deepgram Aura-2 streaming uses a `Flush` message. Its Flush docs warn that very frequent flushes can affect audio quality and cap flushes at 20 per 60 seconds.

Worked example, with illustrative assumptions: an LLM streams 60 tokens per second at about 4 characters per token, or 240 characters per second. If you forward raw tokens to an ElevenLabs socket with the default schedule and no flush, the first generation waits for 120 characters, which is 120 ÷ 240 = 500 ms of buffering before the TTS starts. If instead your orchestrator aggregates to the first sentence ("Sure, I can check that." is 23 characters, about 6 tokens, about 100 ms) and flushes, the TTS starts roughly 400 ms sooner. Frameworks like Pipecat aggregate by sentence for this reason, but custom stacks often do not.

Latency waterfall comparing raw token streaming with a 120-character buffer against sentence aggregation with flush, showing illustrative milliseconds from LLM first token to caller-heard audio

Prosody breaks at chunk joins

Every flush or new request is a seam. If the TTS synthesizes "Your appointment is on" and "Tuesday at 3 PM" as separate requests, pitch resets at the boundary and a pause appears mid-phrase. Callers hear hesitation, and some take the pause as their turn, which triggers an interruption.

To measure it, log the character offsets where text was split, align them to the audio with word timestamps (ElevenLabs returns `alignment`; Deepgram STT returns word times), and compute the silence at each join. Flag joins inside a clause where the pause exceeds a threshold you calibrate on unsplit renders of the same text. Two more streaming gotchas belong in your test plan. ElevenLabs closes an idle WebSocket after 20 seconds by default (configurable up to 180 with `inactivity_timeout`), so a long tool call can force a reconnect and a slower first byte on the next turn. Flush caps can be exhausted on long, sentence-by-sentence explanations. The low-latency TTS guide goes deeper on tuning.

Voice consistency over long calls

A demo clip lasts ten seconds. An intake call can run fifteen minutes and dozens of TTS requests, and several things drift.

  • Speaker identity. Zero-shot and cloned voices can wander from the reference timbre, especially on long or unusual sentences. Measure it the way Seed-TTS-eval does: extract speaker-verification embeddings per turn and compute cosine similarity to an anchor, such as turn one. See evaluating voice cloning for the cloning-specific checks.
  • Decoding failures. Many current TTS models are autoregressive language models over audio codec tokens. The VALL-E 2 paper introduced "Repetition Aware Sampling" specifically to stabilize decoding and avoid an infinite-loop failure. In production that family of failures sounds like repeated words, skipped words, or audio that runs on past the text.
  • Loudness. Requests can come back at different levels, and a quieter turn can be lost on a phone line.
  • Settings drift. Some APIs let you override voice settings per message. ElevenLabs allows `voice_settings` per message on its WebSocket. A code path that sets them differently on one branch changes the voice mid-call.

A cheap detector for decoding failures uses metrics you already collect. LiveKit's `TTSMetrics` exposes `characters_count` and `audio_duration`. Divide characters by seconds for each utterance and compare against that voice's normal speaking rate. A z-score above 3 in either direction means skipped text, repetition, or truncation, and it costs nothing to compute on every call.

For speaker similarity, do not invent a fixed cutoff. Render 200 or so same-voice utterance pairs on clean conditions, record the similarity distribution, and flag production turns below the baseline mean minus three standard deviations.

Tracing TTS in production: an OpenTelemetry span model

A common question is what a good trace model looks like for a cascaded voice agent (ASR to LLM to tools to TTS). Pipecat's OpenTelemetry integration gives a sensible default hierarchy: a `conversation` span, child `turn` spans, and `stt`, `llm` and `tts` spans inside each turn. Turn spans carry `turn.was_interrupted` and `turn.duration_seconds`. TTS spans carry `gen_ai.provider.name`, `gen_ai.request.model`, `voice_id`, `text`, `metrics.character_count` and `metrics.ttfb`.

# Pipecat tracing setup, per Pipecat's OpenTelemetry docs (simplified).
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from pipecat.utils.tracing.setup import setup_tracing
from pipecat.pipeline.worker import PipelineWorker, PipelineParams

setup_tracing(service_name="voice-agent",
              exporter=OTLPSpanExporter(endpoint="http://localhost:4317", insecure=True))

worker = PipelineWorker(
    pipeline,                                   # your existing Pipeline
    params=PipelineParams(enable_metrics=True), # needed for service metrics
    enable_tracing=True,
    enable_turn_tracking=True,
    conversation_id=call_id,
    additional_span_attributes={"tts.voice_config": "aura-2-thalia-en@v3"},
)

The built-in attributes cover service timing and the text sent. They do not tell you whether the caller heard the right thing. Add evaluation spans from an offline scorer that reads the call recording and joins on `conversation.id`. The attribute names below are a suggested convention, not a standard.

AttributeTypeMeaning
`tts.eval.entity_count`intEntities in the text of this turn
`tts.eval.entity_failures`intEntities that failed round-trip exact match
`tts.eval.roundtrip_cer`floatCER of the 8 kHz recording against the span's `text`
`tts.eval.chars_per_sec_z`floatSpeaking-rate z-score versus this voice's baseline
`tts.eval.join_pause_ms_max`intLongest mid-clause pause at a chunk join
`tts.eval.speaker_sim`floatCosine similarity to the call's first-turn embedding
`playout.first_audio_ms`intCaller-heard first audio, from Twilio mark or recording

Because the TTS span stores the exact `text` synthesized, you can score every turn and tie a disputed amount back to the request that produced it. The broader trace design is covered in OpenTelemetry observability for voice agents.

The TTS evaluation scorecard

Copy this scorecard into your test plan. The thresholds are suggested starting points, not industry standards. Set the final numbers from your own baseline, then hold every voice, model and config change to them.

MetricDefinitionMeasured onPassReviewBlock
Entity fidelity rateEntities exactly recovered by both STT engines ÷ total entities8 kHz μ-law≥ 99%97 to 99%< 97%
Round-trip CERCharacter error rate, normalized text, general sentences8 kHz μ-law≤ 2%2 to 5%> 5%
Noisy round-trip CERSame, with listener-side babble at 5 dB SNR8 kHz + noise≤ 2× clean CER2 to 3×> 3×
Caller-heard first audio, TTS shareText ready to first audio on the line, p95Live phone leg≤ 400 ms400 to 600 ms> 600 ms
Join pause anomaliesMid-clause joins over calibrated pause threshold ÷ joinsStreaming renders≤ 2%2 to 5%> 5%
Speaking-rate outliersUtterances with a chars-per-second z-score beyond ±3All renders≤ 0.5%0.5 to 1%> 1%
Speaker similarity driftTurns below baseline mean − 3σ30-turn calls≤ 1%1 to 3%> 3%
Preference vs incumbentAB win rate, 25+ raters, phone audio8 kHz μ-lawLower 95% CI ≥ 45%CI spans 40%Upper CI < 50%

Two rules sit on top of the table. Any single entity failure on a regulated value, such as a dose, an amount owed or a legal disclosure, blocks release regardless of the rate. And an automatic MOS predictor drop of more than about 0.1 on the same voice and scripts triggers a listening review, never a pass or fail on its own.

The test matrix

Cross six content types with five conditions to get 30 cells. With 20 utterances per cell, that is 600 utterances per voice.

Content typeClean 24 kHz8 kHz μ-law8 kHz + noiseToken streamingTurn 30 of a long call
Short confirmations2020202020
Entity-dense (money, dates, phones)2020202020
Names, drugs, product terms2020202020
Long explanations (over 250 characters)2020202020
Questions and lists (prosody)2020202020
Apologies and bad news (tone)2020202020

Size the entity cells deliberately. To claim an entity error rate below 1 percent with 95 percent confidence when you observe zero failures, the "rule of three" says you need at least 3 ÷ 0.01 = 300 entity instances. For a 0.5 percent claim you need 600. If your entity-dense cells hold 100 utterances with three entities each, you are at 300.

How to evaluate a TTS voice for a voice agent

1. Collect real agent text. Pull 500 or more LLM responses from logs or simulations, and add your entity list and domain terms.

2. Normalize in code first. Put a deterministic normalizer between the LLM and TTS for money, dates, phones and codes, with unit tests.

3. Shortlist on fit, not MOS. Use language support, streaming API, price and data terms to cut to 2 to 4 candidates. The best TTS for voice agents roundup is a starting point.

4. Render through your real phone chain. Same transport, resampler and codec you will deploy. Keep the clean renders for comparison.

5. Run round-trip with two STT engines. Score CER, entity fidelity and speaking-rate outliers across the 30-cell matrix.

6. Listen to every failure and every domain term. Confirm whether the TTS or the STT caused each flagged item.

7. Measure streaming in your stack. Capture service TTFB from framework metrics and caller-heard first audio from Twilio marks or recordings, at p50 and p95, with your real buffering and flush policy.

8. Run an AB preference test on phone audio. 25 to 30 raters, counterbalanced order, attention checks, and only the finalists in the pool.

9. Check long-call stability. Run 30-turn calls and score speaker similarity, loudness and rate drift.

10. Gate releases on the scorecard. Re-run the matrix on every voice, model, SDK or normalizer change, and trace production turns with evaluation spans.

For a structured head-to-head between vendors, pair this with a voice agent POC bake-off, or have Evalgent run the matrix across your shortlisted vendors as a neutral party. If you are comparing specific vendors such as xAI and ElevenLabs, our xAI vs ElevenLabs comparison applies this lens to both.

Where independent evaluation fits

Vendor MOS and latency figures are real but measured on vendor terms: their sentences, their sample rate, their definition of first byte. An independent evaluation applies one matrix, phone chain and scorecard to every candidate, so the numbers are comparable.

Evalgent does this at three points: a pre-launch audit of the voice on your scripts and telephony, a bake-off when you are choosing between TTS vendors or models, and regression scoring after each voice or model update. The scorecard above works without us; a third party adds a fixed protocol no vendor tuned and enough calls for tight confidence intervals. For context on how TTS failures show up in the full call, see why voice agents fail in production and our metric thresholds guide.

Frequently asked questions

What is TTS evaluation?

TTS evaluation measures whether a text-to-speech system delivers the right words, with acceptable timing and voice quality, under real conditions. For voice agents that means intelligibility and entity fidelity on 8 kHz phone audio, streaming latency the caller actually hears, consistency across long calls, and paired listener preference. Naturalness scores alone do not cover it.

Why is MOS unreliable for comparing TTS voices?

MOS is a relative score. Cooper and Yamagishi showed one system dropping 1.28 points when re-rated among stronger systems, and Kirkland et al. found identical audio scored 3.21 or 3.53 depending on whether listeners rated naturalness or quality. MOS separates very different systems but struggles with close commercial voices, and numbers from different tests are not comparable.

Can UTMOS or DNSMOS replace human listening tests?

No. In VoiceMOS 2023's zero-shot French TTS track, the best predictor reached a system-level rank correlation of 0.79 and the average was 0.57. DNSMOS was designed for noise suppressors, and UTMOS was trained on 16 kHz data. Use predictors to detect regressions on the same voice and scripts, then confirm with listening.

How many raters does a TTS listening test need?

Plan on 25 to 30 raters. Earlier Blizzard analysis found significance stabilized around 30 listeners. To detect a 60/40 AB preference at 80 percent power you need about 194 independent judgments, roughly doubled to account for each rater's personal taste. Detecting a 55/45 split needs about 783 judgments before that adjustment.

What is round-trip intelligibility testing?

You synthesize text with the TTS, transcribe the audio with an STT model, and score the transcript against the input using CER and exact entity matches. Run it on 8 kHz μ-law audio with two STT vendors. It catches skipped words, wrong numbers and truncation, but STT language models can hide mispronunciations of familiar words.

How do you test TTS pronunciation of drug names?

Build the term list from your formulary, get approved pronunciations from a clinical reviewer, render each term in three carrier sentences, run round-trip with two STT engines, and have a reviewer listen to every term. Fix errors with pronunciation dictionaries or phoneme tags, and re-test after each model change.

What is a good TTS latency for a voice agent?

Measure caller-heard first audio, not vendor model TTFB, which often excludes network and application time. A suggested target is the TTS share at or under 400 ms at p95, from text ready to audio on the line. Check your buffering: forwarding raw tokens into a 120-character buffer can add about 500 ms.

Why does a TTS voice sound worse on phone calls?

Phone legs such as Twilio Media Streams carry 8 kHz μ-law mono audio, which cuts content above about 4 kHz and quantizes quiet detail. Voices that rely on crisp high frequencies lose their edge, and spelled codes suffer most. Resampling quality and transcoding hops add further damage, so evaluate after the phone chain.

The bottom line

TTS evaluation for a voice agent is a test of whether callers hear the right words quickly on a phone line, and MOS answers almost none of that. Measure entity fidelity, round-trip intelligibility, caller-heard latency and long-call stability on 8 kHz audio, use paired preference tests with enough raters for the naturalness question, and gate every voice change on a scorecard you control.

Related Articles