Cascading vs speech-to-speech voice agents: latency budgets, the intelligence tax, cost math, and hybrids

On this page
Most comparisons of cascading vs speech-to-speech voice agents stop at a table: speech-to-speech is faster, cascading is more controllable. That table hides every number you need to make the call.
This guide fills in those numbers. It covers a per-hop latency budget with sourced millisecond ranges. It covers what independent benchmarks say about the reasoning and tool-use gap. It covers cost per minute at stated token assumptions, the hybrid patterns teams now ship, a decision matrix by use case, and a bake-off protocol that treats both architectures fairly.
Evalgent evaluates voice agents independently and has no stake in either architecture. Every number below links to its source or is labeled as an assumption.
Cascading vs speech-to-speech: four architectures, not two
The two-way split is out of date. Production teams now choose between four designs, and the differences between them are about where text exists in the loop.
Cascading voice agent: a pipeline of separate models, speech-to-text (STT), an LLM, and text-to-speech (TTS), plus a turn-taking layer. Text exists at every hand-off.
Speech-to-speech (S2S) voice agent: one audio-native model that takes audio in and produces audio out. Text, if any, is a side output.
Half-cascade: a realtime audio model handles listening and returns text, and a separate TTS speaks it. LiveKit documents this as a supported pipeline type in its pipeline types guide.
S2S with a backend brain: a full-duplex voice model handles the conversation and delegates reasoning and tools to a separate text model. OpenAI's GPT-Live works this way, as our GPT-Live build guide shows.
| Dimension | Cascade | Half-cascade | S2S + backend brain | Pure S2S |
|---|---|---|---|---|
| Where text exists | Every hop | Output only | Backend reasoning | Side transcript only |
| Turn-taking owner | Your VAD/turn model | Realtime model | Voice model | Model |
| Scripted speech (exact words) | Yes | Yes | No | No |
| Live user transcript | Yes, interim | Delayed | Delayed | Delayed |
| Model swap per layer | Yes | TTS only | Backend only | No |
| Prosody-aware listening | No | Yes | Yes | Yes |
| Overlap and backchannels | Bolted on | Model-handled | Model-handled | Model-handled |
LiveKit's own guidance is blunt: "For most production agents, an STT-LLM-TTS pipeline is the right default." Its docs also list GPT-Live as the recommended realtime model for new agents. Both statements can be true, because the right answer depends on which constraint binds first. For the layer-by-layer basics, start with our voice agent stack guide.
How each architecture moves audio
Latency and failure modes follow from the mechanics, so here they are, frame by frame.
Inside a cascade: frames, timers, and streaming hand-offs
A phone call arrives as G.711 audio at 8 kHz, packetized into 20 ms RTP frames. Your media server decodes it and usually resamples to 16 kHz for the STT model. Every resample is a place where audio quality silently degrades, so log the sample rate at each hop.
A voice activity detector (VAD) classifies each frame as speech or not. The VAD feeds an endpointing decision: has the caller finished their turn? In LiveKit Agents, that decision lives under `turn_handling`. The turn handling options reference sets a default `min_delay` of 0.5 seconds and `max_delay` of 3.0 seconds. With the audio turn detector, the defaults drop to 0.3 and 2.5 seconds.
That `min_delay` is dead air you pay on every turn. It is the single largest knob in a cascade's latency budget, and most teams never touch it. Our endpointing guide compares the options.
Deepgram's Flux model fuses transcription and turn detection. It emits `StartOfTurn` and `EndOfTurn` events from the same model that produces the transcript, so the final text is ready the moment the turn ends. Its configuration docs set `eot_threshold` to 0.7 by default and `eot_timeout_ms` to 5000.
Once the transcript is final, the LLM streams tokens. The orchestrator buffers tokens until it has a speakable clause, then sends that clause to TTS. TTS streams audio chunks back, and playback starts before the LLM has finished. LiveKit's sequential pipeline write-up puts a blocking pipeline at 1,000 to 2,000 ms or more and a streaming one at 400 to 800 ms.
Barge-in reverses the flow. When the VAD detects caller speech during agent speech, the orchestrator cancels TTS, flushes queued audio, and records what the caller actually heard. In a cascade you control all of that, which matters later.
The same thing in Pipecat terms
Pipecat models a voice agent as a list of frame processors. A cascaded pipeline reads like transport input, STT, user context aggregator, LLM, TTS, transport output, and assistant context aggregator. Each processor consumes and emits frames, such as audio, transcription, LLM text, and TTS audio frames.
The S2S version replaces the STT, LLM, and TTS processors with one service. Pipecat's OpenAI Realtime service docs describe `OpenAIRealtimeLLMService` as replacing separate STT, LLM, and TTS. Pipecat now runs pipelines inside a `PipelineWorker`, which replaced `PipelineTask`. The practical point is that switching architecture in Pipecat is a few lines. Switching back is just as easy, which is why a fair bake-off is cheap to run.
Inside a speech-to-speech model
The clearest public description of an S2S model is Kyutai's Moshi paper. Moshi models its own speech and the user's speech as two parallel token streams, so there is no explicit speaker turn. It also predicts time-aligned text tokens before audio tokens, a method the authors call Inner Monologue, which improved the linguistic quality of its speech. The paper reports a theoretical latency of 160 ms and 200 ms in practice.
Commercial S2S models differ in detail, but they share three consequences you must design around:
1. The model owns turn-taking. LiveKit's realtime model docs recommend using the model's built-in turn detection, because accurate turn detection needs interim transcripts that realtime models do not provide.
2. Transcripts lag. The same docs warn that user transcriptions "can be considerably delayed and often arrive after the agent's response." Any logic that keys off the transcript runs late.
3. Context and playback can disagree. LiveKit's GPT-Live plugin docs note that GPT-Live does not support message truncation. When the caller interrupts, the model's context still holds its full turn, so it can refer to words the caller never heard.
That third point is a bug class that does not exist in a well-built cascade. It shows up in evaluations as an agent "confirming" details it never said out loud.
The latency budget, hop by hop
Here is a mouth-to-ear budget for a phone call, from the caller's last syllable to the first syllable of the agent's reply. The ranges come from vendor docs where they exist. The rest are labeled assumptions you should replace with your own measurements.

| Hop | What happens | Default cascade (ms) | Tuned cascade (ms) | Basis |
|---|---|---|---|---|
| Inbound transport | 20 ms G.711 frames, carrier, SIP/RTP, jitter buffer | 100 | 80 | Assumption; ITU-T G.114 treats one-way delay up to 150 ms as acceptable for most uses |
| Endpointing | Silence timer or turn model decides the turn ended | 500 | 300 | LiveKit `min_delay` defaults: 0.5 s, or 0.3 s with turn detector |
| STT finalization | Final transcript after endpoint | 150 | 30 | Assumption; fused STT like Flux has text ready at `EndOfTurn` |
| LLM time to first token | Prompt prefill and first token | 600 | 300 | LiveKit cites 300 to 800 ms |
| First-clause buffer | Wait for a speakable phrase | 150 | 80 | Assumption |
| TTS first audio | Synthesis of the first chunk | 200 | 100 | LiveKit cites 100 to 200 ms; ElevenLabs lists Flash at about 75 ms model latency |
| Outbound transport | Network and playout buffer | 100 | 80 | Assumption |
| Total | 1,800 | 970 |
The worked sum for the default column is 100 + 500 + 150 + 600 + 150 + 200 + 100 = 1,800 ms. The tuned column is 80 + 300 + 30 + 300 + 80 + 100 + 80 = 970 ms. Notice that endpointing and LLM prefill account for 61% of the default budget. Those are the two hops to attack first.
What speech-to-speech actually measures
The common claim is that S2S responds in 200 to 300 ms. LiveKit's own comparison table lists 200 to 300 ms for S2S against 300 to 600 ms for a streaming cascade. Independent measurements tell a different story.
The tau-Voice benchmark ran OpenAI's gpt-realtime-1.5, Google's Gemini Live 2.5 Flash native audio, and xAI's Grok voice agent through 278 grounded customer-service tasks. Its latency score averages response latency (user utterance end to agent response) and yield latency (time to stop after an interruption). OpenAI scored fastest at 0.90 s, Google 1.14 s, and xAI 1.15 s.
Artificial Analysis measures time to first audio on Big Bench Audio questions. On its speech-to-speech leaderboard, Grok Voice Think Fast 2.0 High shows 0.70 s, GPT-Realtime-1.5 0.81 s, Gemini 3.8 Live 1.18 s, and GPT-Live-1 (Sol, low) 1.24 s. These are reasoning questions, so thinking time is included.
Add 80 ms of inbound and 80 ms of outbound transport to a 700 to 1,250 ms model-side figure, and a phone-based S2S agent lands around 860 to 1,410 ms. That overlaps the tuned cascade at 970 ms.
For reference, the cross-language study by Stivers and colleagues in PNAS found that the most common gap between a question and its answer falls between 0 and 200 ms in all ten languages studied. No production architecture hits that on first response. Moshi gets close, but it scores 4% on Big Bench Audio on the Artificial Analysis leaderboard. Raon SpeechChat posts 0.04 s time to first audio and 58% on the same benchmark. Speed and reasoning still trade off at the frontier.
The takeaway is not that S2S is slow (see our full-duplex explainer for why overlap is the real prize). It is that S2S's real advantage is not first-response time. Its advantage is overlap handling: backchannels, barge-in, and yielding, which a cascade has to bolt on. Full-Duplex-Bench defines the standard tests for these behaviors: pause handling, backchanneling, turn-taking, and interruption handling. Measure both, separately. Our time to first audio guide defines the measurement, and our turn-taking evaluation guide covers the overlap metrics.
Two latency gotchas that skew comparisons
Eager end-of-turn costs money. Deepgram's Flux launch post says an `eager_eot_threshold` of 0.3 to 0.5 fires `EagerEndOfTurn` 150 to 250 ms earlier than `EndOfTurn`, "at the cost of 50-70% more LLM calls." The same post reports a typical p90 of 1 second and p95 of 1.5 seconds for end-of-turn detection. Budget for the tail, not the median.
Parameter names drift. That launch post calls the silence fallback `eot_silence_threshold_ms`. The current configuration docs call it `eot_timeout_ms`. If your config uses the old name, check that it still takes effect. A silently ignored timeout makes one vendor look worse in a bake-off for reasons that have nothing to do with architecture.
The intelligence tax: what the benchmarks show
"S2S models are dumber" was true in 2024. It is only partly true now, and the part that remains true is the part that matters for agents.

Finding 1: the 2024 reasoning gap was real
When Artificial Analysis released Big Bench Audio, it ran 1,000 spoken questions from Big Bench Hard through several configurations. GPT-4o scored 92% text-to-text. Its native S2S counterpart, GPT-4o Realtime Preview (Oct '24), scored 66%. Text-to-speech landed at 74%, so both audio input and audio output contributed to the drop.
A cascade of Whisper, GPT-4o, and TTS-1 showed "minimal performance degradation compared to pure text processing." That was the original case for cascades: text in the middle preserved the LLM's reasoning.
Finding 2: on short reasoning puzzles, the gap has closed
The same benchmark today shows top S2S models at 96% to 99%. GPT-Realtime-2.1 High scores 96%, Grok Voice Think Fast 2.0 High 97%, and Qwen Audio 3.0 Realtime Plus 99%. If your agent's hardest job is answering one spoken logic question, architecture no longer decides accuracy.
Finding 3: on multi-turn tool tasks, the gap is still large
Tool-heavy work is where the tax lives. tau-Voice extends tau-squared-bench to full-duplex voice, with a simulated caller, real tools, a policy document, and a single valid database end state per task. GPT-5 with reasoning completed 85% of tasks in text. Voice agents reached 31% to 51% under clean audio and 26% to 38% under realistic noise, accents, and turn-taking. They kept only 30% to 45% of text capability.
Three details from the paper matter for builders:
- Most of the loss happens before noise. The text-to-clean-voice drop is the dominant gap for most providers. Realistic audio adds a further 5 points for Google and 12 to 14 points for OpenAI and xAI.
- Accent robustness is provider-specific. xAI lost 38% of its clean capability under accents, while Google was nearly unaffected. Test your caller population, not a generic one.
- The failures are behavioral. The authors attribute 79% to 90% of failures to agent behavior rather than simulator artifacts.
The Artificial Analysis leaderboard shows the same split. Gemini 3.8 Live scores 92% on Big Bench Audio but 30.1% on tau-Voice agentic tasks. GPT-Live-1 (Astra, medium), which delegates reasoning to a text backend, scores 90% and 67.9%. Puzzle accuracy does not predict agent accuracy. The two models at the top of the agentic chart either delegate to a text backend (GPT-Live-1 at 67.9%) or think at length (Gemini 3.8 Live Extended Thinking High at 68.6%).
Finding 4: within-architecture spread beats between-architecture spread
The Artificial Analysis Speech Agent Arena includes "default cascaded systems," each vendor's standard STT, LLM, and TTS stack, scored on task success alongside S2S models.
| System | Architecture | Task success rate | 95% CI |
|---|---|---|---|
| Grok Voice Think Fast 2.0 High | S2S | 94.6% | 91.4 to 96.7 |
| Gemini 3.8 Live | S2S | 93.2% | 89.0 to 95.8 |
| GPT-Live-1 (Sol, low) | S2S + backend | 90.9% | 87.2 to 93.7 |
| ElevenLabs Agents (Scribe v2 Realtime, Gemini 2.5 Flash, Flash v2) | Cascade | 90.5% | 86.7 to 93.4 |
| Cartesia Line (Ink, Gemini 2.5 Flash, Sonic) | Cascade | 77.5% | 68.9 to 84.3 |
| Deepgram Voice Agent (Nova-3, GPT-4o Mini, Aura-2) | Cascade | 73.7% | 68.5 to 78.2 |
| Nova 2.0 Sonic (Mar 2026) | S2S | 57.1% | 51.1 to 62.9 |
S2S task success spans 29.1% to 94.6% on that board. Default cascades span 69.9% to 90.5%. The best default cascade overlaps GPT-Live-1's confidence interval. And the cascades run older, cheaper LLMs: GPT-4o Mini and Gemini 2.5 Flash.
So the "intelligence tax" in 2026 is mostly a model-selection tax. A cascade lets you pay down that tax by swapping in a stronger LLM. An S2S stack makes you wait for the vendor. Our LLM selection guide covers how to pick that model for voice.
Control: what you gain and lose
Control is not one property. It is a bundle of specific capabilities, and each architecture keeps a different subset.
| Capability | Cascade | Half-cascade | S2S + backend | Pure S2S |
|---|---|---|---|---|
| Prompt caching on the reasoning model | Yes | Provider-dependent | Backend only | Provider-dependent |
| Swap the LLM without touching voice | Yes | No | Backend only | No |
| Redact PII before the LLM sees it | Yes | No | No | No |
| Validate tool arguments in code | Yes | Yes | Yes | Yes |
| Speak an exact legal disclosure | Yes | Yes | No | No |
| Truncate context to what was heard | Yes | Provider-dependent | No (GPT-Live) | Provider-dependent |
| Per-hop timestamps and logs | Yes | Partial | Partial | Minimal |
Prompt caching changes the cost curve
In a cascade, the system prompt and tool schemas repeat on every turn. Caching makes that repetition cheap. OpenAI's pricing page lists gpt-6.1-sol at $2.00 per million input tokens and $0.10 cached, a 95% discount. For gpt-realtime-2.1, audio input is $32.00 per million and cached input is $0.40.
The realtime discount matters even more, because a token-billed S2S session re-reads its whole conversation as input on each response. At these rates a cache hit costs about 1% of a miss. Anything that changes the prompt prefix mid-session can forfeit the cached rate for everything after it. Read your provider's caching rules before you add mid-call instruction updates.
PII redaction has to happen somewhere
In a cascade, the transcript passes through your code before the LLM. You can mask card numbers there, or buy redaction from the STT vendor. Deepgram lists streaming redaction at $0.0020 per minute on its pricing page.
In any S2S design, raw caller audio reaches the model provider. Redaction becomes a contract and data-retention question, not a code question. For PCI-scoped flows, many teams route the payment step to DTMF capture or a separate cascade segment.
Deterministic tool layers exist in every architecture
Tool execution runs in your code in all four designs. LiveKit's GPT-Live plugin runs your `@function_tool` methods in your agent process. The difference is who decides to call the tool and with what arguments.
In a cascade, you can inspect the transcript, the LLM's chosen arguments, and the policy state before execution. In GPT-Live's default `responses` delegation, the voice model does not see the tools at all. A backend Responses model chooses them. That split is a feature, but it means tool failures can originate in two models. Our tool call accuracy guide explains how to score each.
Observability is per hop or it is guesswork
A cascade gives you timestamps for VAD start, VAD end, transcript final, first LLM token, first TTS byte, and first played frame. With those, a latency regression points to one component.
An S2S model gives you audio in, audio out, and a late transcript. You can still measure mouth-to-ear latency from a caller-side recording, but you cannot attribute it. If your team needs nightly evaluation runs over every turn, our guide on what to log on every voice agent call lists the fields per architecture.
Cost per minute, worked
Bundled prices have converged, and the spread across architectures is smaller than the spread within them. Here are the current list prices, verified on October 2, 2026.
| Offer | Architecture | Listed price | Source |
|---|---|---|---|
| OpenAI GPT-Live (`gpt-live-1`) | S2S + backend | $0.05/min voice, per second; backend tokens extra | OpenAI pricing |
| xAI Grok Voice | S2S | $0.08/min audio; +$0.01/min on a free xAI number | xAI announcement |
| Deepgram Voice Agent API, Standard | Bundled cascade | $0.075/min pay as you go | Deepgram pricing |
| ElevenLabs Speech Engine | Bundled pipeline | $0.08/min | ElevenLabs API pricing |
| Cartesia Managed Agents | Bundled cascade | $0.06/min, +$0.014/min on a Cartesia number | Cartesia pricing |
A bundled cascade costs about the same as a bundled S2S model. The cost advantage of cascading only appears when you unbundle it and buy each layer yourself.
A do-it-yourself cascade at stated assumptions
Assumptions, all illustrative: a 4-minute call; the STT stream stays open for the full call; the agent speaks about 35% of the time at roughly 900 characters per speaking minute, so about 350 characters per call-minute; three LLM turns per minute; 3,500 input tokens and 70 output tokens per turn; 80% of input tokens served from cache.
| Layer | Choice | Arithmetic per call-minute | Cost |
|---|---|---|---|
| STT | Deepgram Flux English | $0.0065/min promo rate ($0.0077 regular) | $0.0065 |
| LLM | gpt-6.1-sol | 2,100 cache-write tokens at $2.50/M + 8,400 cached at $0.10/M + 210 output at $10/M | $0.0082 |
| TTS | Deepgram Aura-2 | 0.35k characters at $0.030 per 1k | $0.0105 |
| Total | $0.025 |
Swap Aura-2 for ElevenLabs Flash at $0.04 per 1k characters and the total is about $0.029. Swap the LLM for gpt-6-luna and the LLM line falls to under $0.001, for a total near $0.018. Without caching, the gpt-6.1-sol line would be $0.023, so caching alone cuts that line by about two-thirds. Hosting, orchestration, and telephony are excluded because they apply to every architecture.
Token-billed S2S grows with call length
For gpt-realtime-2.1, assume 600 audio tokens per minute of user audio and 1,200 per minute of model audio. That rate matches the per-hour audio prices Artificial Analysis lists for GPT-Realtime-2 ($1.15 input, $4.61 output) at those token prices. Using the same speaking shares, a call-minute generates about 270 new input audio tokens and 420 output audio tokens.
Output costs 420 x $64/M = $0.027. New input costs 270 x $32/M = $0.009. Re-reading the cached context, about 11,600 tokens per minute at minute two, adds about $0.005 at $0.40/M. The total is about $0.040 per minute, and it rises slowly as the context grows. If that context misses the cache, the re-read alone would cost roughly $0.16 per minute instead of $0.005.
Cost per resolved call is the number that matters
Per-minute cost is a distraction if resolution rates differ. Use this formula:
Expected cost per call = C + (1 - r) x H
Here C is the AI cost per call, r is the resolution rate, and H is the cost of the human handling a failed call needs. Assume H = $3.50 for illustration.
For a cascade at C = $0.10 (4 minutes at $0.025) and r = 85%, the expected cost is $0.10 + 0.15 x $3.50 = $0.625. For GPT-Live at C = $0.21 (4 minutes at $0.05, plus an assumed $0.01 of backend tokens) and r = 88%, it is $0.21 + 0.12 x $3.50 = $0.63. They are nearly identical.
The break-even resolution gap is (C_b - C_a) / H. GPT-Live must beat the cascade by (0.21 - 0.10) / 3.50 = 3.1 points to pay for itself. Grok Voice at $0.32 per call needs 6.3 points. Our cost per resolution guide extends this to containment and repeat calls.
That 3-point gap is the uncomfortable part, because detecting it takes a lot of calls. More on that in the bake-off protocol below.
Hybrid patterns that work in production
The most useful architecture decision is often "which hop gets text back." Here are five hybrids, with what each costs you.

1. Half-cascade: audio in, text out
The realtime model listens, hears prosody, and returns text. Your TTS speaks it. You get exact scripted speech, brand voice choice, and stable output. LiveKit notes a quirk this avoids: some realtime models start answering in text only after loading long conversation histories.
The catch is that GPT-Live cannot do this. The LiveKit plugin docs state it "has no text-only response modality." You need a realtime model with a text modality, such as the OpenAI Realtime API.
2. S2S front, text brain behind
GPT-Live's default mode delegates reasoning and tools to a backend Responses model. Its `client` delegation mode goes further: no backend model runs, and your own code answers each delegation with `append_commentary`. That lets you keep your existing LLM, retrieval stack, and policy engine while the voice model handles turn-taking.
Know the limits from the LiveKit docs. Startup history caps at 128 messages and 8,192 tokens. Each append caps at 500 tokens. Context is append-only, so you cannot correct a wrong item. These limits are documented; they are not edge cases.
3. Cascade with S2S-grade turn-taking
Fused STT with end-of-turn detection closes much of the first-response gap while keeping text at every hop. Deepgram reports that Flux "can cut agent response latency by 200-600 ms compared to pipeline approaches." Pair it with eager end-of-turn and speculative LLM calls, and price in the extra LLM calls.
4. Route by turn type
Use S2S for open conversation and a cascade for tool-heavy or regulated steps. LiveKit's pipeline write-up mentions teams exploring this split. Be careful with the hand-off. A GPT-Live agent handoff with different instructions starts a new session, which re-sends history within the 128-message, 8,192-token startup cap. OpenAI's voice cost guide also notes that creating a WebRTC session bills 15 seconds up front, credited once the session runs, so sessions created early or abandoned still cost money.
5. Sidecar STT for the record
Run a separate streaming STT alongside the S2S model to get live captions, compliance transcripts, and turn-detection signals. LiveKit's docs recommend exactly this when you need realtime transcription or its turn detector with a realtime model. You pay for STT twice, at about $0.005 to $0.008 per minute on Deepgram's list prices.
The three configurations in code
The following is simplified. The cascade strings and `TurnHandlingOptions` follow LiveKit's published examples. The half-cascade and GPT-Live forms follow the LiveKit realtime and GPT-Live plugin docs. Check names against your installed version before use.
# Simplified, illustrative. LiveKit Agents (Python).
from livekit.agents import AgentSession, TurnHandlingOptions
from livekit.plugins import openai, silero
# 1) Cascade: text at every hop, tunable endpointing
cascade = AgentSession(
vad=silero.VAD.load(),
stt="deepgram/nova-3:multi",
llm="openai/gpt-4.1-mini",
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
turn_handling=TurnHandlingOptions(
endpointing={"mode": "fixed", "min_delay": 0.3, "max_delay": 2.5},
),
)
# 2) Half-cascade: realtime listening, your TTS speaks
half_cascade = AgentSession(
llm=openai.realtime.RealtimeModel(modalities=["text"]),
tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
)
# 3) S2S with a backend brain: GPT-Live delegates reasoning and tools
s2s_backend = AgentSession(
llm=openai.realtime.GPTLiveModel(
voice="marin",
responses_options={
"model": "gpt-5.6-luna",
"instructions": "Use tools when current information is required.",
},
),
vad=silero.VAD.load(), # needed for playback cut-off on barge-in
)That last `vad` line is a gotcha in its own right. The GPT-Live plugin docs say the session drops its default VAD for this model, and without one, "the agent plays until the model stops on its own."
Decision matrix by use case
Score the constraint that binds first. This matrix is a starting point, built from the capabilities and benchmark findings above. Validate it on your own calls.
| Use case | Lean toward | Why | Must test before launch |
|---|---|---|---|
| Payments, collections, PCI scope | Cascade, or route the payment step to a cascade | Redaction before the LLM, exact disclosures | Disclosure verbatim rate, redaction misses |
| Healthcare intake and scheduling | Cascade or half-cascade | Exact wording, live transcript, audit trail | Entity accuracy on names, dates, medications |
| Tier-1 support with many tools | S2S + backend, or cascade with a strong LLM | tau-Voice shows tool tasks drive the gap | Database end-state success under noise |
| Outbound qualification | Cascade with fused turn detection | Cost at volume, scripted openers | Answering machine handling, opener timing |
| Coaching, companionship, sales role-play | Pure S2S or S2S + backend | Prosody, overlap, backchannels | Interruption and backchannel handling |
| Multilingual, accented caller base | Test both; no default | Accent robustness varied by provider in tau-Voice | Task success per accent cohort |
| Noisy environments (drive-thru, field) | Cascade with noise suppression | You can tune VAD and STT per environment | Success at 10 dB and 5 dB SNR |
| Web or app voice with good audio | S2S or half-cascade | Clean wideband audio, no PSTN leg | Latency p90, overlap quality |
Use a weighted score when the call is close. Assign weights to five criteria: task success, p90 latency, overlap handling, control requirements, and cost per resolved call. Weights should sum to 100. A regulated contact center might use 35, 15, 10, 25, and 15. A consumer companion app might use 20, 20, 35, 5, and 20.
How to run a fair cascade vs speech-to-speech bake-off
Most architecture comparisons are unfair by accident. They run the S2S model over WebRTC and the cascade over the phone, or they compare a vendor dashboard's latency against a stopwatch. Use this protocol to compare like with like. Our guide to comparing voice agents on the same test cases covers the scenario design in more depth.
1. Freeze a scenario set. Write 40 scenarios that cover your top intents, two policy edge cases per intent, and at least five multi-tool tasks. Give each scenario a single correct end state, as tau-Voice does, so success is checkable without a judge.
2. Cross scenarios with conditions. Run each scenario under five conditions: clean audio, 10 dB SNR babble noise, 5 dB SNR street noise, two accent cohorts drawn from your caller base, and a scripted interruption at a fixed point. That gives 200 scenario-condition cells.
3. Use the same transport leg for both. Route both architectures through the same SIP trunk at 8 kHz G.711. A 16 kHz WebRTC path flatters whichever system gets it. See our SIP vs WebRTC guide for why.
4. Use the same tools and the same mock backend. Point both agents at identical tool schemas and a seeded test database. Score success by comparing the final database state to the expected one.
5. Measure at the caller side. Record both channels on the caller's leg. Compute response latency as the gap from caller speech end to agent speech start, using one VAD for both systems. Report p50 and p90, not means.
6. Score overlap separately. Measure agent interruption rate, yield latency after a barge-in, and selectivity on backchannels such as "mm-hmm." These are the metrics where S2S should win, so give it the chance.
7. Size the sample before you start. To detect 80% vs 85% task success at 80% power and alpha 0.05, you need about 903 calls per arm. To detect 85% vs 88%, the 3-point break-even from the cost section, you need about 2,033 per arm.
8. Run three trials per cell. Artificial Analysis averages three trials where available on its tau-Voice charts. Single runs hide variance from model sampling and simulator randomness.
9. Report cost per resolved call with confidence intervals. Use the formula above, with your real human handling cost. If the intervals overlap, the architectures tie on your data, and control requirements should decide.
Here is the sample-size and latency arithmetic in Python, so you can rerun it with your own rates:
# Two-proportion sample size and caller-side latency percentiles (illustrative)
from math import ceil
from statistics import quantiles
Z_ALPHA, Z_BETA = 1.96, 0.8416 # two-sided alpha 0.05, power 0.80
def calls_per_arm(p1: float, p2: float) -> int:
var = p1 * (1 - p1) + p2 * (1 - p2)
return ceil((Z_ALPHA + Z_BETA) ** 2 * var / (p1 - p2) ** 2)
print(calls_per_arm(0.80, 0.85)) # 903 (about 902.6)
print(calls_per_arm(0.85, 0.88)) # 2033
def latency_report(caller_end_ms: list[float], agent_start_ms: list[float]) -> dict:
gaps = [a - c for c, a in zip(caller_end_ms, agent_start_ms)]
pct = quantiles(gaps, n=100)
return {"p50_ms": pct[49], "p90_ms": pct[89], "turns": len(gaps)}Those sample sizes explain why so many architecture decisions rest on anecdotes. A 50-call pilot cannot separate two systems that differ by 3 to 5 points. This is where an independent evaluator helps: Evalgent runs the same scenario matrix against both stacks on the same transport leg and reports the intervals, so the decision rests on your calls rather than a leaderboard. For the statistics behind it, see our A/B testing guide for voice agents.
Non-obvious gotchas from production
These are the failure modes that show up after launch, not in demos.
- S2S agents confirm things the caller never heard. Without truncation, the model's context holds the full interrupted turn. Test it by interrupting mid-confirmation and checking what the agent claims it said.
- Late transcripts break transcript-keyed logic. Escalation rules, sentiment triggers, and keyword guards that read the user transcript fire after the model has already answered. Move those checks to tool calls or a sidecar STT.
- Scripted speech is not available on full-duplex models. LiveKit's docs say GPT-Live "can't speak a script word for word." Even with a TTS attached for `session.say()`, the model can talk over it.
- History seeding has hard caps. GPT-Live accepts 128 messages and 8,192 tokens at startup, oldest dropped first. Long callbacks with prior context lose the oldest facts silently.
- Eager turn detection inflates LLM spend. Deepgram puts it at 50% to 70% more LLM calls. Price it into the cascade's cost line, not just the latency line.
- Default cascades on leaderboards use older LLMs. The arena's cascaded entries run GPT-4o Mini and Gemini 2.5 Flash. A cascade with a current model is a different system.
For a broader test plan for the S2S side, see our guides to testing speech-to-speech voice agents and evaluating a speech-to-speech model. For framework-level latency differences, see LiveKit vs Pipecat latency, and for xAI's stack specifically, our Grok voice agent guide.
Frequently asked questions
Is speech-to-speech faster than a cascaded voice pipeline?
Usually, but by less than vendors claim. Independent measurements put leading S2S models at about 0.7 to 1.25 seconds model-side. A default cascade with a 500 ms endpointing delay lands near 1.8 seconds on a phone call. A tuned cascade with fused turn detection lands near 1 second. S2S's clearer advantage is handling overlap, interruptions, and backchannels.
Do speech-to-speech models reason worse than text models?
On short spoken puzzles, not anymore. Top S2S models now score 96% to 99% on Big Bench Audio, up from 66% for GPT-4o Realtime in 2024. On multi-turn tool tasks, the gap remains. In tau-Voice, voice agents kept only 30% to 45% of a text model's task success under realistic audio.
What is a half-cascade voice agent?
A half-cascade pairs a realtime audio model, which listens and returns text, with a separate TTS that speaks. You keep prosody-aware listening and gain exact scripted speech and voice choice. It requires a realtime model with a text-only response modality. GPT-Live does not offer one, according to LiveKit's plugin docs.
How much does a speech-to-speech voice agent cost per minute?
GPT-Live lists $0.05 per minute of voice time, billed per second, plus backend tokens. Grok Voice lists $0.08 per minute, plus $0.01 on a free xAI number. Token-billed gpt-realtime-2.1 works out near $0.04 per minute at typical speaking shares. A self-assembled cascade can run $0.018 to $0.029 per minute.
Which is better for regulated industries, cascading or speech-to-speech?
Cascading, in most cases. It lets you redact PII before the LLM sees it, speak exact disclosures, and keep a text trail at every hop. In any S2S design, raw audio reaches the model provider and scripted speech is limited. Many teams use S2S for open conversation and route regulated steps to a cascade.
Can I use GPT-Live with my own LLM?
Yes, through client delegation. In that mode no backend model runs. Your code receives each delegation event, works out the answer with your own LLM or database, and returns it with `append_commentary`. Function tools are ignored in this mode, so your code owns all reasoning and tool execution.
How many test calls do I need to compare two architectures?
More than most pilots run. Detecting 80% vs 85% task success at 80% power needs about 900 calls per arm. Detecting a 3-point gap, such as 85% vs 88%, needs about 2,000 per arm. Smaller tests can catch large failures but cannot rank systems that differ by a few points.
Should I start with cascading or speech-to-speech in 2026?
Start with a cascade if you need exact wording, PII redaction, or per-layer debugging. Start with S2S, ideally with a text backend, if overlap, prosody, and natural barge-in drive your outcomes. In both cases, keep the scenario suite architecture-neutral. Then you can switch later and measure the difference instead of guessing.
The bottom line
Speech-to-speech models have closed the reasoning gap on short questions and win on overlap and prosody, but tuned cascades now match them within a few hundred milliseconds on first response, cost less when unbundled, and still lead on control and multi-step tool work. Pick by the constraint that binds first, then prove the choice with a same-scenario, same-transport bake-off sized to detect the resolution gap that would pay for the more expensive option.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: how to detect regressions from model updates you don't control
Voice agent regression testing for changes you don't control: pin each layer, run a nightly pinned-vs-latest canary, and use paired stats to detect drops.
Read more