Evalgent
Back to Blog
Voice AI Evaluation

Speech-to-Text (STT) Providers for Voice Agents: 2026 Comparison

Deepesh Jayal
14 min read
Speech-to-Text (STT) Providers for Voice Agents: 2026 Comparison

Speech recognition is the front door of a voice agent. Everything downstream — the model's reasoning, the tools it calls, the answer it speaks — runs on whatever the STT layer thought it heard. Get it wrong and a misheard digit or name quietly breaks the call. This guide compares the fifteen STT providers most relevant to voice agents in 2026, with the numbers as each vendor publishes them, and an honest warning about why those numbers don't compare the way a table implies.

If you want the decision framework rather than the roster, our how to choose an STT guide covers the criteria; this post is the named comparison that sits beside it.

How to read this comparison

One rule matters more than any cell in the table: you cannot rank STT providers on accuracy across their published numbers. Each vendor reports word error rate on a different benchmark — one on clean read speech, another on call-center audio, another on conversational or Indic data. A 3% figure on read audiobooks and a 12% figure on noisy phone calls are not the same test, so lining them up as if they were is exactly the vendor-metrics trap an evaluator has to avoid. We name the benchmark next to every accuracy figure for that reason.

The same caution applies to latency and price. Latency figures are vendor claims under favorable conditions; callers feel delay quickly, and the ITU-T G.114 standard puts the comfortable one-way limit at 150ms. Pricing changes often and varies by tier, so everything below is "as published, September 2026" — confirm current numbers on each vendor's page, and see our WER vs CER guide on reading accuracy figures.

The 15 STT providers compared

Figures below are as published (September 2026). "WER" always names its benchmark, because the benchmarks differ and the numbers are not directly comparable.

Provider (flagship model)Streaming latencyAccuracy (WER · benchmark)LanguagesDiarizationOn-premPrice (as published)Compliance
Deepgram (Nova-3 / Flux)<300ms~5.3% batch (Deepgram real-world set)50+YesYes~$0.0077/min streamSOC 2, HIPAA, GDPR
AssemblyAI (Universal-3.5)~130ms sync~7.0% (Pipecat, realtime)BroadYesYes~$0.15–0.45/hrSOC 2, HIPAA
OpenAI (gpt-4o-transcribe / Whisper)Tunable (realtime)~2.5% (FLEURS/LibriSpeech, clean)~99 (Whisper)No (external)Whisper only~$0.003–0.017/min*SOC 2, HIPAA (API)
ElevenLabs (Scribe v2 Realtime)<150ms~3.1% (FLEURS)~99Yes (32 spk)No~$0.39/hrSOC 2, ISO 27001, PCI, HIPAA
Google Cloud (Chirp 3)Not publishedNot published85+Yes (subset)No~$0.016/minSOC 2, HIPAA, GDPR
Microsoft Azure (AI Speech)Sub-secondNot published140+ localesYes (35 spk)Yes (containers)~$0.0167/min RTSOC, ISO, HIPAA, PCI, FedRAMP
Amazon TranscribeNot published~12.9% (AssemblyAI 2026 set)100+ / 54 streamYesNo~$0.024/min batchHIPAA, SOC, PCI, FedRAMP
Speechmatics (Ursa)Sub-1s~1.1% (Pipecat, Aug 2026)55+Yes (free)Yes (+on-device)Usage-basedISO 27001, SOC 2, HIPAA
NVIDIA (Riva / Parakeet)~160ms–2s (configurable)~6.3% (Parakeet v3, Open ASR Leaderboard)~25 (Parakeet)YesYes (core model)License + GPU (not per-min)You own the environment
Sarvam (Saaras v4)<150ms TTFT~19.3% (IndicVoices, 10 langs)22 IndianYesYes (+air-gapped)~₹30/hrSOC 2, ISO 27001, India DPDP
Gladia (Solaria)~103ms partialTops Switchboard (no single WER)100+YesYes~$0.75/hr RTSOC 2, HIPAA, GDPR, ISO
Soniox (v4)Ultra-low (not published)~1.25% English (own 60-lang study)60+Yes (bundled)In-region~$0.12/hr RTSOC 2, ISO 27001, HIPAA
Rev AI (Reverb)~500msVendor "low WER" (not named)58 / 9 streamYesNot confirmed~$0.20/hrSOC 2, HIPAA, GDPR, PCI
IBM Watson STTNot publishedNot published~38 modelsYes (2-way)Yes (containers)~$0.01/minSOC 2, HIPAA, GDPR, ISO
Cartesia (Ink-2 / Ink-Whisper)~100ms~8% (AppTek multi-accent)English (Ink-2)No (turn detection)Yes~$0.009/min (est.)SOC 2, HIPAA, GDPR, PCI

\*OpenAI's gpt-4o-transcribe is token-billed, not per-minute — model it separately.

Provider notes

Deepgram — Nova-3 for transcription and Flux for native turn detection make it a purpose-built voice-agent choice, with sub-300ms streaming and true self-hosting. Best for low-latency agents at scale that want on-prem control.

AssemblyAI — Accuracy-first, with rich audio intelligence and easy self-serve HIPAA. Its sync API returns a finished transcript in about 130ms, and self-hosted parity is a plus for regulated English workloads.

OpenAI — Flexible and cheap if you are already in the ecosystem, with open-source Whisper for free self-hosting. The catch: no native diarization, and the newest models are token-billed API-only.

ElevenLabs (Scribe) — A strong newer entrant: sub-150ms realtime, ~99 languages, diarization up to 32 speakers, and deep compliance including PCI and data residency. Cloud-only, no self-host.

Google Cloud (Chirp 3) — Broad language coverage and data-residency controls, best for teams already on GCP. Google publishes little committed latency or WER, so verify on your audio.

Microsoft Azure — Turnkey enterprise compliance (including FedRAMP), 140+ locales, and container deployment for partial on-prem. A safe default if you are Azure-native.

Amazon Transcribe — Managed and tightly integrated with AWS, with strong compliance. Independent tests put its WER higher than the specialists, so weigh accuracy against ecosystem convenience.

Speechmatics — Accuracy and accent robustness are its calling card, with free diarization and flexible deployment down to on-device. Strong for global, regulated, multi-accent traffic.

NVIDIA (Riva) — Structurally different: self-hosted software on your GPUs, not a per-minute API, so its "price" is licensing plus infrastructure. Its Parakeet models sit on the public Open ASR Leaderboard. Best when data sovereignty or GPU-optimized scale matters most.

Sarvam — The Indic specialist: 22 Indian languages, code-mixing, India data residency, and air-gapped options. The obvious pick for Indian-language agents; its WER is on the hard IndicVoices benchmark, not a clean-English one.

Gladia — Widest language coverage with strong diarization and EU/US residency, built streaming-first at ~103ms partials. A good fit for multilingual contact centers.

Soniox — The value leader on paper, with bundled diarization and translation at the lowest per-hour rate and heavy multilingual coverage. In-region deployment rather than full on-prem.

Rev AI — English-first, with a human-transcription fallback and solid compliance including a 99.99% uptime SLA. Streaming language coverage is narrow (nine languages).

IBM Watson — Built for regulated enterprises that want container/OpenShift on-prem deployment and custom-model tuning. Publishes little public accuracy data, so plan to measure it yourself.

Cartesia (Ink) — Real-time-first with ~100ms latency and native turn detection, tuned for structured data like phone numbers and dates. Note two things: Ink-2 is English-only, and it offers turn detection rather than speaker diarization.

How to actually choose

The table narrows the field; it does not pick the winner. Start from your constraints: real-time streaming is non-negotiable for a voice agent, so drop batch-only options first. Then weight the criteria to your use case — entity accuracy and compliance dominate for finance and healthcare, language coverage for global or Indic traffic, and self-hosting for data sovereignty. For regulated work, map compliance to a framework like the NIST AI Risk Management Framework, and treat the STT as one layer of the wider voice agent stack. Model cost per resolved call, not per minute — accuracy that reduces escalations is worth more than a fractional rate difference, the discipline our vendor evaluation guide applies.

Then do the one thing no roundup can do for you: run your finalists on your own calls. Because the published WER figures sit on different benchmarks, and none of them are your callers, the only trustworthy comparison is the same calls, scored the same way, across every candidate — the benchmark-on-your-own-data principle, applied to the STT layer with our STT evaluation methodology.

Comparing STT providers with Evalgent

Evalgent turns this table into a decision grounded in your traffic. Scenarios drive real calls through each candidate STT, and Profiles vary accent, pace, and line noise, so you measure accuracy on the audio your callers actually produce — not a vendor's clean benchmark. Metrics score word error rate weighted for your critical entities, latency at the tail, and downstream task success, with one fixed definition applied to every provider so the comparison is finally apples to apples. Evaluations run the suite at concurrency, and Reviews let your team replay any call to hear where a provider mishears. Because Evalgent has no stake in which STT you pick, the ranking it produces reflects your calls, not a leaderboard.

The result is an STT choice you can defend: measured on your own audio, scored identically, and re-checked when a provider ships a model update. To compare the STT providers on your shortlist against your own calls, book a demo.

The bottom line

There is no single best STT for voice agents — there is the best one for your accents, your entities, your compliance needs, and your budget. The fifteen providers here differ most on language coverage, on-prem availability, and how their accuracy holds up on real phone audio versus clean benchmarks.

Use the table to shortlist, weight the criteria to your use case, and then benchmark the finalists on your own calls. Vendor numbers tell you what each provider wanted to show you; your own measurement tells you which one will actually hear your callers.

Frequently asked questions

What is the best STT for voice agents in 2026?

There is no universal best — it depends on your accents, domain entities, languages, compliance needs, and budget. Purpose-built low-latency options like Deepgram, ElevenLabs Scribe, Gladia, Soniox, and Cartesia target real-time agents; specialists like Sarvam lead for Indian languages. Shortlist from the comparison, then benchmark your finalists on your own calls, since vendor accuracy numbers use different benchmarks.

Which STT has the lowest latency for voice agents?

Several providers publish roughly 100–150ms streaming or first-token latency in 2026 — Cartesia around 100ms, Gladia around 103ms for partials, ElevenLabs Scribe under 150ms, and Deepgram under 300ms end to end. But these are vendor claims under favorable conditions. Measure latency at the tail (p90/p95) on your own audio and network, because real-world latency differs from a benchmark figure.

Can you compare STT providers on their published WER?

No, not directly. Each vendor reports word error rate on a different benchmark — clean read speech, call-center audio, conversational or Indic data — so the numbers measure different tests. A low WER on audiobook audio says nothing about noisy phone calls. Always name the benchmark, and treat published WER as directional only; the sole comparable measure is running each provider on your own calls.

Which STT providers offer on-prem or self-hosted deployment?

As published in 2026, Deepgram, AssemblyAI, Speechmatics, NVIDIA (Riva is self-hosted by design), Sarvam, IBM Watson (containers), and Cartesia offer on-prem or self-hosted options; Microsoft Azure offers container deployment. ElevenLabs, Google, and Amazon are cloud-only with data-residency controls instead. Confirm current availability with each vendor, since self-hosting terms and parity change.

How much do STT APIs cost for voice agents?

Streaming STT commonly runs from roughly $0.002 to $0.017 per minute in 2026, depending on provider and tier, with batch cheaper than streaming. Two gotchas: OpenAI's newest models are token-billed rather than per-minute, and NVIDIA is licensing plus GPU cost, not a per-minute API. Prices change often, so confirm on each vendor's pricing page and model cost per resolved call, not per minute.

Which STT is best for multilingual or non-English voice agents?

For breadth, providers citing 60–100+ languages — Gladia, Soniox, ElevenLabs Scribe, OpenAI Whisper, and Amazon (batch) — cover the most ground. For Indian languages and code-mixed speech, Sarvam is the specialist with 22 Indian languages and India data residency. Match the provider to your specific languages and verify accuracy on real audio in those languages, not on an aggregate benchmark.

Do STT providers support HIPAA and SOC 2 for voice agents?

Most enterprise STT providers publish SOC 2 Type II, and many offer HIPAA with a BAA, plus GDPR and often ISO 27001 or PCI — including Deepgram, AssemblyAI, ElevenLabs, Azure, AWS, Google, Speechmatics, Gladia, Soniox, Rev AI, and IBM. Self-hosted options like NVIDIA put compliance in your hands. Confirm the exact certifications and whether they cover the specific STT service on each vendor's trust page.

How should you choose an STT provider for a voice agent?

Drop batch-only providers first, since voice agents need streaming. Weight the criteria to your use case — entity accuracy and compliance for regulated work, language coverage for global traffic, self-hosting for data sovereignty. Shortlist two or three, then run them on your own calls with one fixed scoring method. Decide on that measured result, not the roundup.

Related Articles