Evalgent
Back to Blog
Voice AI Evaluation

Text-to-Speech (TTS) Providers for Voice Agents: 2026 Comparison

Deepesh Jayal
14 min read
Text-to-Speech (TTS) Providers for Voice Agents: 2026 Comparison

Text-to-speech is the voice your caller actually hears. It decides whether the agent sounds human or robotic, whether it says a dollar amount clearly, and — through its latency — whether the reply feels instant or laggy. This guide compares the thirteen live TTS providers most relevant to voice agents in 2026, with the numbers as each vendor publishes them, and an honest note on why "naturalness" is the one column you cannot rank from a table.

For the decision framework rather than the roster, our how to choose a TTS guide covers the criteria and the low-latency TTS deep-dive covers time-to-first-audio; this is the named comparison beside them.

How to read this comparison

Three cautions before the table. First, naturalness does not compare across vendors. Every provider reports quality on a different listener study or arena — one crowdsourced Elo, one blind A/B, one internal preference test — so a mean opinion score or "ranked #1" claim describes that test, not a shared scale defined by a standard like ITU-T P.800. Judge speech synthesis quality by ear, on your own scripts.

Second, latency figures use different definitions — time to first byte versus time to first audio, single-stream versus concurrent — and are vendor best-cases. Callers feel the pause, and the ITU-T G.114 standard puts the comfortable one-way limit at 150ms, so measure it yourself. Third, pricing units differ — per character, per minute, or credits — so normalize before comparing. Everything below is "as published, September 2026"; confirm current numbers on each vendor's page.

The 13 TTS providers compared

Figures are as published (September 2026). Latency mixes time-to-first-audio and time-to-first-byte, as noted; naturalness is intentionally omitted from the table because vendor scores are not comparable.

Provider (flagship model)Streaming latencyVoices / languagesVoice cloningEmotion controlOn-premPrice (as published)Compliance
ElevenLabs (Eleven v3 / Flash v2.5)~75ms (Flash)10,000+ voices / 70+ langsYes (instant + pro)Yes (audio tags)No (VPC + residency)$0.05–0.10 / 1k charsSOC 2, ISO, PCI, HIPAA
Cartesia (Sonic-3.6)~40–90ms (vendor)40+ langsYesYesYes (air-gapped)~$0.03 / minSOC 2, HIPAA, PCI, GDPR
Deepgram (Aura-2)~90ms50 voices / 7 langsNoLimitedYes~$0.03 / 1k charsSOC 2, HIPAA, GDPR
Rime (Arcana v3 / Mist v2)~120–225ms300+ voices / ~10 langsNot documentedYes (paralinguistics)Yes$0.03–0.05 / 1k charsSOC 2, HIPAA
Hume (Octave 2)~100–200ms11 langsYesYes (strongest)No$0.05–0.15 / 1k charsSOC 2, GDPR, HIPAA (ent.)
OpenAI (gpt-4o-mini-tts)~300–600ms13 voices / multilingualNoYes (prompt)No~$0.015 / min*SOC 2, GDPR (HIPAA via Azure)
Google Cloud (Chirp 3: HD)Low (not published)380+ voices / 75+ langsYes (instant, gated)Yes (SSML)No~$30 / 1M charsSOC 2, HIPAA, GDPR
Microsoft Azure (Neural HD / CNV)Low (not published)600+ voices / 150+ langsYes (gated)Yes (SSML styles)Yes (containers)$16–48 / 1M charsSOC, ISO, HIPAA, PCI, FedRAMP
Amazon Polly (Generative)Streaming (not published)100+ voices / 40+ langsYes (Brand Voice, gated)Yes (SSML)No$16–30 / 1M charsHIPAA, SOC, PCI, FedRAMP
Sarvam (Bulbul v3)<250ms TTFB~30 voices / 11 IndianNot confirmedYesYes (air-gapped)~₹30 / 10k charsSOC 2, ISO, India DPDP
Inworld (TTS-2 / Flash)~25–100ms TTFB200+ langsYesYes (steerable)Yes$5–15 / 1M charsSOC 2, HIPAA, GDPR
NVIDIA (Magpie / Riva)~32–79ms~12 langsZero-shot conditioningPartialYes (core model)License + GPUYou own the environment
Resemble (Chatterbox)~500ms~23 langsYes (instant + pro)Yes (exaggeration)Yes (air-gapped)~$0.0005 / secSOC 2, HIPAA, GDPR, C2PA

\*OpenAI's TTS is token-billed; the per-minute figure is an estimate — model it separately.

Provider notes

ElevenLabs — The widest voice library and the expressiveness leader, with Flash v2.5 at roughly 75ms for realtime and deep compliance. Cloud-only, but VPC and data residency cover most enterprise needs.

Cartesia (Sonic) — Built for the fastest realtime agents, with a state-space architecture, strong emotional control, and true self-hosting down to air-gapped. Vendor latency is a best case; measure it on your own load.

Deepgram (Aura-2) — Cost-efficient enterprise TTS tuned to speak numbers, dates, and passwords clearly, with self-hosting for regulated work. No voice cloning and a fixed voice set are the trade-offs.

Rime — Hyper-realistic through paralinguistics — breaths, disfluencies, vocal fry — with on-prem support and low on-device latency. A strong fit for support and healthcare agents that must sound human.

Hume (Octave 2) — The expressiveness and emotion leader, controllable with natural-language acting instructions rather than SSML. Cloud-only, and full compliance sits on the enterprise tier.

OpenAI — Simple and steerable if you are already in the ecosystem, with tone controlled by plain instructions. The catches: no voice cloning, higher latency, and HIPAA only via Azure OpenAI.

Google Cloud (Chirp 3: HD) — Broad language coverage, instant custom voice, and expressive HD voices, best for teams on GCP. Cloning is gated to allow-listed customers, and Google publishes little committed latency.

Microsoft Azure — The enterprise deployment story: on-prem containers, FedRAMP, 150+ languages, and controlled professional voice cloning. Custom Neural Voice is gated behind a use-case review.

Amazon Polly — Managed and AWS-native with a new generative engine and bidirectional streaming, plus broad compliance. Custom "Brand Voice" is a gated engagement, not self-serve cloning.

Sarvam (Bulbul) — The Indic specialist: 11 Indian languages, expressive output, India data residency, and air-gapped deployment. The obvious pick for Indian-language agents needing sovereignty.

Inworld — A standout low-latency and cost story — roughly 25–100ms first byte, from about $5 per million characters — with steerable emotion and full self-hosting. Strong for expressive realtime agents at scale.

NVIDIA (Magpie / Riva) — Self-hosted software on your GPUs, not a per-minute API, with 32–79ms first audio and open weights. Best when data sovereignty or GPU-native scale matters most; compliance is on you.

Resemble (Chatterbox) — Compliance- and provenance-focused, with C2PA watermarking and air-gapped deployment for deepfake-sensitive use. Higher latency (~500ms) makes it better for quality-first than snappy realtime.

Recently discontinued — do not build on these

Two names that still rank in older roundups are no longer live in 2026, so treat them as dead:

  • PlayHT / PlayAI — the team was acqui-hired and the platform sunset at the end of 2025, with the API and voice clones removed.
  • LMNT — shut down around September 2026; its integrations have been removed from major frameworks.

If you are on either, migrate to a live provider above. The point of a 2026 comparison is to keep you off a dependency that just disappeared — the same lock-in risk our vendor lock-in guide warns about.

How to actually choose

The table narrows the field; your ear and your use case pick the winner. Real-time streaming is non-negotiable for a voice agent, so weight latency heavily and confirm it under load, not from a best-case number. Then match the rest to your needs — expressiveness and prosody) for brand voices, language coverage for global or Indic traffic, self-hosting for data sovereignty, and pronunciation accuracy on the numbers and names your calls depend on. For regulated work, map compliance to a framework like the NIST AI Risk Management Framework rather than a vendor's certification badge alone.

Then do the one thing a roundup can't: listen to your finalists on your own scripts and calls. Because naturalness scores sit on different studies, and latency figures are best cases, the only trustworthy comparison is the same lines spoken by each provider, judged by your team and your callers — the benchmark-on-your-own-data principle applied to the voice layer, and part of the wider voice agent stack decision.

Comparing TTS providers with Evalgent

Evalgent turns this table into a decision grounded in your calls. Scenarios run your real scripts through each candidate TTS inside full conversations, and Profiles vary the caller and line conditions, so you hear how each voice lands on the audio your callers actually get. Metrics capture time to first audio at the tail and pronunciation accuracy on your critical entities — names, numbers, dates — with one fixed definition per provider, and Reviews let your team replay and rate the voices side by side to settle naturalness by ear rather than by a vendor's arena score. Evaluations run at concurrency, so a latency claim is tested under load. Because Evalgent has no stake in which TTS you pick, the comparison reflects your calls, not a leaderboard.

The result is a voice choice you can defend: heard on your own scripts, measured for latency and pronunciation, and re-checked when a provider ships a new model. To compare the TTS providers on your shortlist against your own calls, book a demo.

The bottom line

There is no single best TTS for voice agents — there is the best one for your latency budget, your languages, your brand voice, and your compliance needs. The thirteen live providers here differ most on realtime latency, voice cloning, on-prem availability, and how expressive they sound on your actual scripts.

Use the table to shortlist, weight latency and pronunciation for your use case, and then listen to the finalists on your own calls. Vendor naturalness scores tell you which test a provider won; your own ears tell you which voice your callers will trust.

Frequently asked questions

What is the best TTS for voice agents in 2026?

There is no universal best — it depends on your latency budget, languages, brand voice, and compliance needs. Realtime-first options like ElevenLabs Flash, Cartesia, Inworld, and NVIDIA lead on latency; Hume and Rime lead on expressiveness; Sarvam owns Indian languages. Shortlist from the comparison, then judge the finalists by ear on your own scripts, since naturalness scores use different studies.

Which TTS has the lowest latency for voice agents?

Several providers publish very low time-to-first-audio in 2026 — Inworld around 25–100ms, NVIDIA Magpie around 32–79ms, ElevenLabs Flash around 75ms, Cartesia around 40–90ms, and Hume near 100ms in instant mode. But these are vendor best cases using different definitions. Measure time to first audio at the tail, under your expected concurrency, on your own network before trusting a figure.

Can you compare TTS providers on naturalness or MOS?

Not directly. Each vendor reports quality on a different listener study — a crowdsourced arena, a blind A/B test, an internal preference panel — so the numbers describe different tests, not a shared scale. A "ranked #1" or MOS claim is real for that study only. The only comparable measure is your own team and callers rating the voices on your scripts.

Which TTS providers offer voice cloning?

As published in 2026, ElevenLabs, Cartesia, Hume, Inworld, and Resemble offer instant or professional voice cloning; Google, Azure, and Amazon offer cloning that is gated behind approval (Custom Voice or Brand Voice). Deepgram and OpenAI do not clone voices. Confirm current terms and any use-case review with each vendor, since cloning access and safeguards change.

Which TTS providers support on-prem or self-hosting?

Cartesia, Deepgram, Rime, Sarvam, Inworld, NVIDIA (self-hosted by design), Resemble, and Microsoft Azure (containers) offer on-prem or self-hosted options as published in 2026. ElevenLabs, Hume, OpenAI, Google, and Amazon are cloud-only, though several offer VPC or data-residency controls. Verify current availability and parity with each vendor before committing.

How much does TTS cost for voice agents?

Pricing units differ, which makes headline numbers misleading: per character (ElevenLabs, Deepgram, Rime, Hume, Google, Azure, Amazon), per minute or credits (Cartesia), per second (Resemble), token-billed (OpenAI), or license plus GPU (NVIDIA). Character rates commonly run from a few dollars to about $30 per million. Normalize to cost per minute of speech, and confirm current pricing on each vendor's page.

Which TTS is best for multilingual or Indian-language voice agents?

For breadth, Inworld (200+ languages), Azure (150+), Google (75+), and ElevenLabs (70+) cover the most ground. For Indian languages and code-mixed speech, Sarvam is the specialist with 11 Indian languages and India data residency. Match the provider to your exact languages and judge pronunciation and naturalness by ear in those languages, not on an aggregate score.

Are PlayHT and LMNT still available for voice agents?

No. PlayHT/PlayAI was acqui-hired and its platform sunset at the end of 2025, with the API and voice clones removed. LMNT shut down around September 2026, and its integrations have been pulled from major frameworks. If you depend on either, migrate to a live provider from the comparison above, and treat the episode as a reminder to keep your voice layer swappable.

Related Articles