Evalgent
Back to Blog
Voice AI Evaluation

Multilingual TTS for Voice Agents: How to Choose

Deepesh Jayal
12 min read
Multilingual TTS for Voice Agents: How to Choose

# Multilingual TTS for voice agents: how to choose

Quick answer

Multilingual TTS for voice agents is text-to-speech that sounds native in every language it speaks, not one English voice reading Spanish. Choose it by judging pronunciation and prosody per language, code-switching, and name and number handling by ear, on your own scripts, in each language you serve.

Your vendor demos a warm, natural English voice. Then a Spanish caller hears the same voice stumble over "gracias," flatten every question into a statement, and read a phone number in English digits. The voice was never multilingual. It was an English model doing an impression of Spanish.

This guide is about the output side of a multilingual voice agent: the speech your callers hear. It covers what native quality means per language. It covers whether to use one voice or many, how code-switching breaks pronunciation, and how names and numbers go wrong. It closes with a method for choosing and testing multilingual TTS on your own scripts.

For the input side, how the agent hears your callers, see our companion post on multilingual accuracy for voice agents. This post is its text-to-speech counterpart.

Why an English voice reading Spanish is not multilingual TTS

Speech synthesis turns text into an audio waveform. A multilingual engine claims to do this well in more than one language. The claim is easy to make and hard to verify, because a demo picks the flattering case.

> Multilingual TTS: text-to-speech that produces native-sounding pronunciation and prosody in each language it supports, rather than applying one language's sound system to another's text. It is judged per language, not by a single overall voice rating.

Many engines are strongest in English and weaker everywhere else. They may apply English pronunciation rules to Spanish text, use English intonation on a Spanish question, or mangle accented characters. To an English speaker the clip sounds fine. To a native Spanish speaker it sounds foreign, and it signals that your brand did not take their language seriously.

The fix is the same discipline we apply across evaluation. You judge each language on its own, on scripts that match your real calls, using native ears. A single "voice quality" rating averages away the exact failures your multilingual callers will hear. For the broader method, see our overview of voice agent evaluation.

Where multilingual TTS quality actually breaks

Several distinct problems hide under a single voice rating. Each has a different cause and a different test. Naming them is the difference between a real evaluation and a demo that only proves the English case. The table below is the core scope for judging multilingual TTS.

Evaluation dimensionWhy it is hardHow to judge it
Per-language pronunciationEngines map text to sound using rules that skew toward well-resourced languages; rare sounds get approximatedHave a native speaker rate pronunciation on your own script for each language, flagging every mispronounced word
Per-language prosodyIntonation, stress, and rhythm differ by language; a question in one language is flat in anotherScore naturalness by ear on statements, questions, and lists, comparing against how a native speaker would say them
Native-speaker naturalnessAccent and voice quality can sound foreign even when words are correctCollect mean opinion scores from native listeners per language, not one blended rating
Code-switchingA single sentence mixes languages; many engines assume one language per utteranceFeed real mixed-language lines and listen for wrong-language pronunciation at the switch point
Names, numbers, currencyDigits, dates, and money are read by language-specific rules; foreign names have no ruleTest a fixed list of names, phone numbers, dates, and amounts in each language and check every reading
One voice across languagesKeeping one identity consistent while switching sound systems is technically hardCompare the same voice across languages for timbre drift and per-language quality gaps

Every row is scored on your scripts, by native speakers, not on a vendor reel. The rest of this guide walks through the rows that trip teams up most.

Native-quality pronunciation and prosody per language

Pronunciation is the foundation. If the engine says the words wrong, nothing downstream saves the call. Correct pronunciation means mapping text to the right phonemes, the distinct sound units of a language. Each language has its own inventory, and some sounds in one language do not exist in another.

Engines that lean on English tend to approximate foreign sounds with the nearest English one. The result is an accent a native speaker hears immediately. A useful reference point is the International Phonetic Alphabet, which gives every sound a symbol. When a provider supports pronunciation control, they usually accept phonetic input or SSML tags to force a specific reading.

Prosody is the second half, and it is where "correct but foreign" lives. Prosody) covers intonation, stress, and rhythm, the music of speech above the individual sounds. Languages carry meaning differently through prosody. A rising pitch that marks a question in one language may be neutral in another. Stress can fall on different syllables and change the word.

An engine that gets phonemes right but prosody wrong sounds like a robot reading a foreign phrasebook. It is intelligible and still off-putting. To judge it, listen to the same script across sentence types: a plain statement, a yes-or-no question, a list of options, and a number read aloud. A native speaker will tell you within seconds whether the melody is right. We go deeper on this in our guide to pronunciation for voice agents.

Naturalness, and how to measure it by ear

Naturalness is whether the voice sounds like a person from that language community. It combines pronunciation, prosody, and voice quality into one impression. There is no automatic score that captures it reliably, so you measure it with human listeners.

The standard instrument is the mean opinion score, or MOS. Listeners rate clips on a scale, usually one to five, and you average the ratings. The critical rule for multilingual work is to collect MOS per language, from native speakers of that language. An English speaker cannot rate whether a Spanish voice sounds native. A pooled MOS across languages hides your weakest one, exactly as a blended accuracy score does on the input side.

Set the test up honestly. Use your own scripts, not the provider's showcase sentences. Include the hard cases: long numbers, proper names, and questions. Recruit native listeners per language. Report each language separately with its own score and its own pass bar. That is the difference between knowing a voice sounds native and hoping it does.

One multilingual voice or separate voices per language

A common design question is whether to use a single voice that speaks every language or a distinct voice per language. There is no universal answer. The right choice depends on your brand, your languages, and how often a single call crosses languages.

A single multilingual voice keeps one identity across the whole service. That consistency is valuable for brand, and it is essential when one call switches languages mid-stream, because a voice swap mid-conversation is jarring. The risk is uneven quality. The same voice may sound native in English and slightly foreign in another language, since keeping one timbre while changing sound systems is hard.

Separate per-language voices let you pick the best-sounding voice in each language independently. You optimize each language on its own. The cost is a loss of a single consistent identity. You now maintain and evaluate several voices instead of one. If your provider builds voices by cloning, the per-language route also raises quality and consent questions we cover in evaluating voice cloning for voice agents.

Use a single multilingual voice when brand consistency matters most and calls frequently cross languages. Use separate voices when each language stands alone and per-language quality is the priority. Whichever you choose, judge every language the same way: native listeners, your scripts, a pass bar per language.

Code-switching pronunciation within a sentence

Code-switching is alternating languages within one conversation or one sentence. It is normal speech for many bilingual US callers, and your agent has to produce it, not just understand it. Say a caller's name or a product name is Spanish. An English agent still has to say it correctly inside an English sentence.

Many engines assume one language per utterance. Give them a mixed line and they apply one language's rules to the whole thing. The foreign word inside gets an English pronunciation, or the switch point produces an audible stumble. The sentence "Your cita is confirmed for Thursday" should say "cita" the Spanish way, not rhyme it with English.

Test this directly with real mixed lines from your scripts. Include foreign proper names, place names, and product names embedded in the base language. Listen specifically at the switch point. SSML phoneme tags can force a correct reading for known problem words. Check whether your provider supports them and whether they actually fix the case. The accent side of this problem is covered in our guide to accent handling for voice agents.

Names, numbers, currency, and foreign words

The most common real-world failures are the least glamorous ones. Voice agents read a lot of names, phone numbers, dates, times, currency amounts, and addresses. Every one of these is read by language-specific rules. A multilingual engine has to switch rule sets with the language.

Numbers are read differently across languages, and digit grouping, decimal marks, and currency phrasing all change. An amount like $1,250.50 has a specific spoken form in English and a different one in Spanish. Dates and times reorder. Phone numbers get grouped and paced by convention. An engine that reads a Spanish amount with English number words is wrong even though every digit is present.

Proper names are the hardest case because there is no single rule. A name from a third language, embedded in either base language, has to be pronounced plausibly. Text normalization, converting written forms like "$1,250.50" or "Dr." into spoken words, is where much of this happens under the hood, and it is language-dependent. Reading a written form correctly depends on mapping each grapheme, or written character, to the right sound in the active language.

Build a fixed test list and reuse it for every provider and every language. Include real customer names from your data, phone numbers, dates, times, currency amounts, and any product or place names your agent says often. Have a native speaker check every reading. This single list will surface more failures than any generic demo script.

Latency and cost across languages

Quality is not the only axis. A voice agent has a real-time budget, and synthesis latency eats into it. Measure time to first audio and streaming behavior for each language and voice, because performance can differ across them. A voice that sounds great but starts slowly still hurts the call.

Cost also varies. Providers price synthesis by characters or by audio duration. Some languages produce longer audio for the same meaning, which changes the per-call cost. Higher-quality or cloned voices can carry a premium. When you compare options, hold quality, latency, and cost in view together rather than optimizing one in isolation. For running these comparisons on your real traffic, see benchmarking voice agents on your own data.

How to choose and test multilingual TTS

Here is a repeatable method. It works whether you are choosing a provider or auditing one already in production. The principle throughout is per-language judgment by native ears on your own scripts.

1. List every language and mix you serve. Pull your call logs and find the real distribution. Include code-switching as its own category. Note the names, numbers, and currency formats that appear most. This list defines the test scope.

2. Build one fixed test script per language. Include statements, questions, lists, long numbers, dates, currency amounts, and real proper names. Add mixed-language lines for code-switching. Reuse this script across every provider so comparisons are fair.

3. Synthesize the script on each candidate. Generate audio for the same script from every provider and voice you are considering. Keep the settings comparable. Save the clips so native listeners can score them blind.

4. Score naturalness per language with native listeners. Collect mean opinion scores from native speakers of each language separately. Never pool across languages. Flag every mispronounced word and every unnatural intonation pattern.

5. Check names, numbers, and currency explicitly. Walk your fixed list and confirm each reading in each language. A single wrong number reading is a failing case, not a rounding error.

6. Test code-switching at the switch point. Play the mixed lines and listen for wrong-language pronunciation where the languages meet. Confirm whether SSML or phoneme controls fix the known problem words.

7. Measure latency and cost per language. Record time to first audio and the per-call cost for each language and voice. A voice that fails your real-time budget is out, however natural it sounds.

8. Set a pass bar per language and re-test on a schedule. Define the minimum acceptable naturalness and the required pronunciation accuracy for each language. Providers update models, so re-run the script on fresh candidates to catch regressions before callers do.

This method is the multilingual case of a general vendor evaluation. For choosing against explicit criteria, see how to evaluate voice agent vendors.

How Evalgent judges per-language TTS quality

Evalgent is an independent, third-party evaluator for AI voice agents. We do not build the engines we test, so we have no stake in which one wins. That independence is the point. A provider that picks its own languages, scripts, and listeners can produce almost any headline number. An outside measurement on your real scripts, by native ears, cannot be gamed the same way.

For multilingual TTS, we score each language separately on your own scripts. We collect naturalness ratings from native speakers per language. We walk a fixed list of names, numbers, dates, and currency amounts, test code-switching at the switch point, and measure latency and cost per language and voice. You get a per-language scorecard with pass bars, not a single voice rating that flatters your strongest language. For the case for using an outside evaluator at all, see independent voice AI evaluation.

Frequently asked questions

What is multilingual TTS for voice agents?

Multilingual TTS for voice agents is text-to-speech that produces native-sounding speech in each language the agent serves. It is judged per language, not by one overall voice rating. A strong evaluation scores pronunciation, prosody, and naturalness separately for every language, plus code-switching and the reading of names, numbers, and currency, on your own scripts.

How do you tell if a TTS voice is truly native in a language?

Have native speakers of that language rate the voice on your own scripts, not the provider's demo lines. Listen for correct phonemes, natural intonation on questions and lists, and plausible names and numbers. Collect a mean opinion score per language from native listeners. If only English speakers judged it, the rating tells you nothing about the other languages.

Should you use one multilingual voice or separate voices per language?

Use one multilingual voice when brand consistency matters most and calls frequently switch languages mid-conversation, since a voice swap mid-call is jarring. Use separate per-language voices when each language stands alone and you want the best-sounding voice in each. Either way, evaluate every language the same way: native listeners, your scripts, and a pass bar per language.

Can text-to-speech handle code-switching within a single sentence?

Some engines can, many cannot. Many assume one language per utterance, so a foreign word inside a sentence gets the wrong pronunciation or produces an audible stumble at the switch point. Test it with real mixed lines from your scripts and listen where the languages meet. SSML phoneme tags can force correct readings for known problem words if the provider supports them.

How should a voice agent pronounce names, numbers, and currency in each language?

Each language reads digits, dates, times, and currency by its own rules, so the engine must switch rule sets with the language. An amount like $1,250.50 has a distinct spoken form per language. Build a fixed list of real names, phone numbers, dates, and amounts, then have a native speaker confirm every reading in each language you serve.

How do you test multilingual TTS on your own scripts?

Build one fixed script per language with statements, questions, lists, long numbers, dates, currency, real names, and mixed-language lines. Synthesize it on every candidate voice with comparable settings. Have native listeners score naturalness per language, walk the names and numbers explicitly, and check code-switching at the switch point. Reuse the same script across providers so comparisons stay fair.

Does multilingual TTS add latency or cost per language?

It can. Synthesis latency, measured as time to first audio, may differ across languages and voices, and a slow voice hurts the call even when it sounds natural. Cost varies too, since some languages produce longer audio for the same meaning and higher-quality voices carry a premium. Measure latency and cost per language alongside quality, not in isolation.

What is the difference between multilingual TTS and multilingual STT evaluation?

Multilingual TTS evaluation judges the output the caller hears: pronunciation, prosody, and naturalness of synthesized speech per language. Multilingual STT evaluation judges the input the agent hears: how accurately it transcribes each language. Both need per-language testing on real data, but they measure opposite ends of the call. Our companion post covers the input, or STT, side.

The bottom line

Multilingual TTS is only real if it sounds native in every language, not just English. Judge pronunciation, prosody, code-switching, and name and number handling per language, by native ears, on your own scripts.

Want to hear how your agent sounds in every language you serve? Book a demo and we will show you where the voice passes and where it falls short.

Related Articles