Evalgent
Back to Blog
Voice AI Evaluation

Cascading vs speech-to-speech voice agents: latency budgets, the intelligence tax, cost math, and hybrids

Deepesh Jayal
Updated
25 min read
Cascading vs speech-to-speech voice agents: latency budgets, the intelligence tax, cost math, and hybrids
On this page

Most comparisons of cascading vs speech-to-speech voice agents stop at a table: speech-to-speech is faster, cascading is more controllable. That table hides every number you need to make the call.

This guide fills in those numbers. It covers a per-hop latency budget with sourced millisecond ranges. It covers what independent benchmarks say about the reasoning and tool-use gap. It covers cost per minute at stated token assumptions, the hybrid patterns teams now ship, a decision matrix by use case, and a bake-off protocol that treats both architectures fairly.

Evalgent evaluates voice agents independently and has no stake in either architecture. Every number below links to its source or is labeled as an assumption.

92% to 66%
GPT-4o text vs native speech-to-speech on Big Bench Audio, Dec 2024 (Artificial Analysis)
30-45%
Share of text task success that full-duplex voice agents retained in tau-Voice (arXiv 2603.13686)
0.90-1.15 s
Measured latency scores for OpenAI, Google, and xAI voice agents in tau-Voice
$0.05/min
GPT-Live voice price, billed per second, backend tokens extra (OpenAI)

Cascading vs speech-to-speech: four architectures, not two

The two-way split is out of date. Production teams now choose between four designs, and the differences between them are about where text exists in the loop.

Cascading voice agent: a pipeline of separate models, speech-to-text (STT), an LLM, and text-to-speech (TTS), plus a turn-taking layer. Text exists at every hand-off.

Speech-to-speech (S2S) voice agent: one audio-native model that takes audio in and produces audio out. Text, if any, is a side output.

Half-cascade: a realtime audio model handles listening and returns text, and a separate TTS speaks it. LiveKit documents this as a supported pipeline type in its pipeline types guide.

S2S with a backend brain: a full-duplex voice model handles the conversation and delegates reasoning and tools to a separate text model. OpenAI's GPT-Live works this way, as our GPT-Live build guide shows.

DimensionCascadeHalf-cascadeS2S + backend brainPure S2S
Where text existsEvery hopOutput onlyBackend reasoningSide transcript only
Turn-taking ownerYour VAD/turn modelRealtime modelVoice modelModel
Scripted speech (exact words)YesYesNoNo
Live user transcriptYes, interimDelayedDelayedDelayed
Model swap per layerYesTTS onlyBackend onlyNo
Prosody-aware listeningNoYesYesYes
Overlap and backchannelsBolted onModel-handledModel-handledModel-handled

LiveKit's own guidance is blunt: "For most production agents, an STT-LLM-TTS pipeline is the right default." Its docs also list GPT-Live as the recommended realtime model for new agents. Both statements can be true, because the right answer depends on which constraint binds first. For the layer-by-layer basics, start with our voice agent stack guide.

How each architecture moves audio

Latency and failure modes follow from the mechanics, so here they are, frame by frame.

Inside a cascade: frames, timers, and streaming hand-offs

A phone call arrives as G.711 audio at 8 kHz, packetized into 20 ms RTP frames. Your media server decodes it and usually resamples to 16 kHz for the STT model. Every resample is a place where audio quality silently degrades, so log the sample rate at each hop.

A voice activity detector (VAD) classifies each frame as speech or not. The VAD feeds an endpointing decision: has the caller finished their turn? In LiveKit Agents, that decision lives under `turn_handling`. The turn handling options reference sets a default `min_delay` of 0.5 seconds and `max_delay` of 3.0 seconds. With the audio turn detector, the defaults drop to 0.3 and 2.5 seconds.

That `min_delay` is dead air you pay on every turn. It is the single largest knob in a cascade's latency budget, and most teams never touch it. Our endpointing guide compares the options.

Deepgram's Flux model fuses transcription and turn detection. It emits `StartOfTurn` and `EndOfTurn` events from the same model that produces the transcript, so the final text is ready the moment the turn ends. Its configuration docs set `eot_threshold` to 0.7 by default and `eot_timeout_ms` to 5000.

Once the transcript is final, the LLM streams tokens. The orchestrator buffers tokens until it has a speakable clause, then sends that clause to TTS. TTS streams audio chunks back, and playback starts before the LLM has finished. LiveKit's sequential pipeline write-up puts a blocking pipeline at 1,000 to 2,000 ms or more and a streaming one at 400 to 800 ms.

Barge-in reverses the flow. When the VAD detects caller speech during agent speech, the orchestrator cancels TTS, flushes queued audio, and records what the caller actually heard. In a cascade you control all of that, which matters later.

The same thing in Pipecat terms

Pipecat models a voice agent as a list of frame processors. A cascaded pipeline reads like transport input, STT, user context aggregator, LLM, TTS, transport output, and assistant context aggregator. Each processor consumes and emits frames, such as audio, transcription, LLM text, and TTS audio frames.

The S2S version replaces the STT, LLM, and TTS processors with one service. Pipecat's OpenAI Realtime service docs describe `OpenAIRealtimeLLMService` as replacing separate STT, LLM, and TTS. Pipecat now runs pipelines inside a `PipelineWorker`, which replaced `PipelineTask`. The practical point is that switching architecture in Pipecat is a few lines. Switching back is just as easy, which is why a fair bake-off is cheap to run.

Inside a speech-to-speech model

The clearest public description of an S2S model is Kyutai's Moshi paper. Moshi models its own speech and the user's speech as two parallel token streams, so there is no explicit speaker turn. It also predicts time-aligned text tokens before audio tokens, a method the authors call Inner Monologue, which improved the linguistic quality of its speech. The paper reports a theoretical latency of 160 ms and 200 ms in practice.

Commercial S2S models differ in detail, but they share three consequences you must design around:

1. The model owns turn-taking. LiveKit's realtime model docs recommend using the model's built-in turn detection, because accurate turn detection needs interim transcripts that realtime models do not provide.

2. Transcripts lag. The same docs warn that user transcriptions "can be considerably delayed and often arrive after the agent's response." Any logic that keys off the transcript runs late.

3. Context and playback can disagree. LiveKit's GPT-Live plugin docs note that GPT-Live does not support message truncation. When the caller interrupts, the model's context still holds its full turn, so it can refer to words the caller never heard.

That third point is a bug class that does not exist in a well-built cascade. It shows up in evaluations as an agent "confirming" details it never said out loud.

The latency budget, hop by hop

Here is a mouth-to-ear budget for a phone call, from the caller's last syllable to the first syllable of the agent's reply. The ranges come from vendor docs where they exist. The rest are labeled assumptions you should replace with your own measurements.

Latency waterfall for a phone call: default cascade about 1,800 ms, tuned cascade about 970 ms, speech-to-speech 860 to 1,410 ms, split by network, endpointing, STT, LLM, and TTS
HopWhat happensDefault cascade (ms)Tuned cascade (ms)Basis
Inbound transport20 ms G.711 frames, carrier, SIP/RTP, jitter buffer10080Assumption; ITU-T G.114 treats one-way delay up to 150 ms as acceptable for most uses
EndpointingSilence timer or turn model decides the turn ended500300LiveKit `min_delay` defaults: 0.5 s, or 0.3 s with turn detector
STT finalizationFinal transcript after endpoint15030Assumption; fused STT like Flux has text ready at `EndOfTurn`
LLM time to first tokenPrompt prefill and first token600300LiveKit cites 300 to 800 ms
First-clause bufferWait for a speakable phrase15080Assumption
TTS first audioSynthesis of the first chunk200100LiveKit cites 100 to 200 ms; ElevenLabs lists Flash at about 75 ms model latency
Outbound transportNetwork and playout buffer10080Assumption
Total1,800970

The worked sum for the default column is 100 + 500 + 150 + 600 + 150 + 200 + 100 = 1,800 ms. The tuned column is 80 + 300 + 30 + 300 + 80 + 100 + 80 = 970 ms. Notice that endpointing and LLM prefill account for 61% of the default budget. Those are the two hops to attack first.

What speech-to-speech actually measures

The common claim is that S2S responds in 200 to 300 ms. LiveKit's own comparison table lists 200 to 300 ms for S2S against 300 to 600 ms for a streaming cascade. Independent measurements tell a different story.

The tau-Voice benchmark ran OpenAI's gpt-realtime-1.5, Google's Gemini Live 2.5 Flash native audio, and xAI's Grok voice agent through 278 grounded customer-service tasks. Its latency score averages response latency (user utterance end to agent response) and yield latency (time to stop after an interruption). OpenAI scored fastest at 0.90 s, Google 1.14 s, and xAI 1.15 s.

Artificial Analysis measures time to first audio on Big Bench Audio questions. On its speech-to-speech leaderboard, Grok Voice Think Fast 2.0 High shows 0.70 s, GPT-Realtime-1.5 0.81 s, Gemini 3.8 Live 1.18 s, and GPT-Live-1 (Sol, low) 1.24 s. These are reasoning questions, so thinking time is included.

Add 80 ms of inbound and 80 ms of outbound transport to a 700 to 1,250 ms model-side figure, and a phone-based S2S agent lands around 860 to 1,410 ms. That overlaps the tuned cascade at 970 ms.

For reference, the cross-language study by Stivers and colleagues in PNAS found that the most common gap between a question and its answer falls between 0 and 200 ms in all ten languages studied. No production architecture hits that on first response. Moshi gets close, but it scores 4% on Big Bench Audio on the Artificial Analysis leaderboard. Raon SpeechChat posts 0.04 s time to first audio and 58% on the same benchmark. Speed and reasoning still trade off at the frontier.

The takeaway is not that S2S is slow (see our full-duplex explainer for why overlap is the real prize). It is that S2S's real advantage is not first-response time. Its advantage is overlap handling: backchannels, barge-in, and yielding, which a cascade has to bolt on. Full-Duplex-Bench defines the standard tests for these behaviors: pause handling, backchanneling, turn-taking, and interruption handling. Measure both, separately. Our time to first audio guide defines the measurement, and our turn-taking evaluation guide covers the overlap metrics.

Two latency gotchas that skew comparisons

Eager end-of-turn costs money. Deepgram's Flux launch post says an `eager_eot_threshold` of 0.3 to 0.5 fires `EagerEndOfTurn` 150 to 250 ms earlier than `EndOfTurn`, "at the cost of 50-70% more LLM calls." The same post reports a typical p90 of 1 second and p95 of 1.5 seconds for end-of-turn detection. Budget for the tail, not the median.

Parameter names drift. That launch post calls the silence fallback `eot_silence_threshold_ms`. The current configuration docs call it `eot_timeout_ms`. If your config uses the old name, check that it still takes effect. A silently ignored timeout makes one vendor look worse in a bake-off for reasons that have nothing to do with architecture.

The intelligence tax: what the benchmarks show

"S2S models are dumber" was true in 2024. It is only partly true now, and the part that remains true is the part that matters for agents.

Bar chart of the speech intelligence gap: GPT-4o at 92% text versus 66% speech-to-speech on Big Bench Audio, and tau-Voice agents at 26 to 51% versus 85% text

Finding 1: the 2024 reasoning gap was real

When Artificial Analysis released Big Bench Audio, it ran 1,000 spoken questions from Big Bench Hard through several configurations. GPT-4o scored 92% text-to-text. Its native S2S counterpart, GPT-4o Realtime Preview (Oct '24), scored 66%. Text-to-speech landed at 74%, so both audio input and audio output contributed to the drop.

A cascade of Whisper, GPT-4o, and TTS-1 showed "minimal performance degradation compared to pure text processing." That was the original case for cascades: text in the middle preserved the LLM's reasoning.

Finding 2: on short reasoning puzzles, the gap has closed

The same benchmark today shows top S2S models at 96% to 99%. GPT-Realtime-2.1 High scores 96%, Grok Voice Think Fast 2.0 High 97%, and Qwen Audio 3.0 Realtime Plus 99%. If your agent's hardest job is answering one spoken logic question, architecture no longer decides accuracy.

Finding 3: on multi-turn tool tasks, the gap is still large

Tool-heavy work is where the tax lives. tau-Voice extends tau-squared-bench to full-duplex voice, with a simulated caller, real tools, a policy document, and a single valid database end state per task. GPT-5 with reasoning completed 85% of tasks in text. Voice agents reached 31% to 51% under clean audio and 26% to 38% under realistic noise, accents, and turn-taking. They kept only 30% to 45% of text capability.

Three details from the paper matter for builders:

  • Most of the loss happens before noise. The text-to-clean-voice drop is the dominant gap for most providers. Realistic audio adds a further 5 points for Google and 12 to 14 points for OpenAI and xAI.
  • Accent robustness is provider-specific. xAI lost 38% of its clean capability under accents, while Google was nearly unaffected. Test your caller population, not a generic one.
  • The failures are behavioral. The authors attribute 79% to 90% of failures to agent behavior rather than simulator artifacts.

The Artificial Analysis leaderboard shows the same split. Gemini 3.8 Live scores 92% on Big Bench Audio but 30.1% on tau-Voice agentic tasks. GPT-Live-1 (Astra, medium), which delegates reasoning to a text backend, scores 90% and 67.9%. Puzzle accuracy does not predict agent accuracy. The two models at the top of the agentic chart either delegate to a text backend (GPT-Live-1 at 67.9%) or think at length (Gemini 3.8 Live Extended Thinking High at 68.6%).

Finding 4: within-architecture spread beats between-architecture spread

The Artificial Analysis Speech Agent Arena includes "default cascaded systems," each vendor's standard STT, LLM, and TTS stack, scored on task success alongside S2S models.

SystemArchitectureTask success rate95% CI
Grok Voice Think Fast 2.0 HighS2S94.6%91.4 to 96.7
Gemini 3.8 LiveS2S93.2%89.0 to 95.8
GPT-Live-1 (Sol, low)S2S + backend90.9%87.2 to 93.7
ElevenLabs Agents (Scribe v2 Realtime, Gemini 2.5 Flash, Flash v2)Cascade90.5%86.7 to 93.4
Cartesia Line (Ink, Gemini 2.5 Flash, Sonic)Cascade77.5%68.9 to 84.3
Deepgram Voice Agent (Nova-3, GPT-4o Mini, Aura-2)Cascade73.7%68.5 to 78.2
Nova 2.0 Sonic (Mar 2026)S2S57.1%51.1 to 62.9

S2S task success spans 29.1% to 94.6% on that board. Default cascades span 69.9% to 90.5%. The best default cascade overlaps GPT-Live-1's confidence interval. And the cascades run older, cheaper LLMs: GPT-4o Mini and Gemini 2.5 Flash.

So the "intelligence tax" in 2026 is mostly a model-selection tax. A cascade lets you pay down that tax by swapping in a stronger LLM. An S2S stack makes you wait for the vendor. Our LLM selection guide covers how to pick that model for voice.

Control: what you gain and lose

Control is not one property. It is a bundle of specific capabilities, and each architecture keeps a different subset.

CapabilityCascadeHalf-cascadeS2S + backendPure S2S
Prompt caching on the reasoning modelYesProvider-dependentBackend onlyProvider-dependent
Swap the LLM without touching voiceYesNoBackend onlyNo
Redact PII before the LLM sees itYesNoNoNo
Validate tool arguments in codeYesYesYesYes
Speak an exact legal disclosureYesYesNoNo
Truncate context to what was heardYesProvider-dependentNo (GPT-Live)Provider-dependent
Per-hop timestamps and logsYesPartialPartialMinimal

Prompt caching changes the cost curve

In a cascade, the system prompt and tool schemas repeat on every turn. Caching makes that repetition cheap. OpenAI's pricing page lists gpt-6.1-sol at $2.00 per million input tokens and $0.10 cached, a 95% discount. For gpt-realtime-2.1, audio input is $32.00 per million and cached input is $0.40.

The realtime discount matters even more, because a token-billed S2S session re-reads its whole conversation as input on each response. At these rates a cache hit costs about 1% of a miss. Anything that changes the prompt prefix mid-session can forfeit the cached rate for everything after it. Read your provider's caching rules before you add mid-call instruction updates.

PII redaction has to happen somewhere

In a cascade, the transcript passes through your code before the LLM. You can mask card numbers there, or buy redaction from the STT vendor. Deepgram lists streaming redaction at $0.0020 per minute on its pricing page.

In any S2S design, raw caller audio reaches the model provider. Redaction becomes a contract and data-retention question, not a code question. For PCI-scoped flows, many teams route the payment step to DTMF capture or a separate cascade segment.

Deterministic tool layers exist in every architecture

Tool execution runs in your code in all four designs. LiveKit's GPT-Live plugin runs your `@function_tool` methods in your agent process. The difference is who decides to call the tool and with what arguments.

In a cascade, you can inspect the transcript, the LLM's chosen arguments, and the policy state before execution. In GPT-Live's default `responses` delegation, the voice model does not see the tools at all. A backend Responses model chooses them. That split is a feature, but it means tool failures can originate in two models. Our tool call accuracy guide explains how to score each.

Observability is per hop or it is guesswork

A cascade gives you timestamps for VAD start, VAD end, transcript final, first LLM token, first TTS byte, and first played frame. With those, a latency regression points to one component.

An S2S model gives you audio in, audio out, and a late transcript. You can still measure mouth-to-ear latency from a caller-side recording, but you cannot attribute it. If your team needs nightly evaluation runs over every turn, our guide on what to log on every voice agent call lists the fields per architecture.

Cost per minute, worked

Bundled prices have converged, and the spread across architectures is smaller than the spread within them. Here are the current list prices, verified on October 2, 2026.

OfferArchitectureListed priceSource
OpenAI GPT-Live (`gpt-live-1`)S2S + backend$0.05/min voice, per second; backend tokens extraOpenAI pricing
xAI Grok VoiceS2S$0.08/min audio; +$0.01/min on a free xAI numberxAI announcement
Deepgram Voice Agent API, StandardBundled cascade$0.075/min pay as you goDeepgram pricing
ElevenLabs Speech EngineBundled pipeline$0.08/minElevenLabs API pricing
Cartesia Managed AgentsBundled cascade$0.06/min, +$0.014/min on a Cartesia numberCartesia pricing

A bundled cascade costs about the same as a bundled S2S model. The cost advantage of cascading only appears when you unbundle it and buy each layer yourself.

A do-it-yourself cascade at stated assumptions

Assumptions, all illustrative: a 4-minute call; the STT stream stays open for the full call; the agent speaks about 35% of the time at roughly 900 characters per speaking minute, so about 350 characters per call-minute; three LLM turns per minute; 3,500 input tokens and 70 output tokens per turn; 80% of input tokens served from cache.

LayerChoiceArithmetic per call-minuteCost
STTDeepgram Flux English$0.0065/min promo rate ($0.0077 regular)$0.0065
LLMgpt-6.1-sol2,100 cache-write tokens at $2.50/M + 8,400 cached at $0.10/M + 210 output at $10/M$0.0082
TTSDeepgram Aura-20.35k characters at $0.030 per 1k$0.0105
Total$0.025

Swap Aura-2 for ElevenLabs Flash at $0.04 per 1k characters and the total is about $0.029. Swap the LLM for gpt-6-luna and the LLM line falls to under $0.001, for a total near $0.018. Without caching, the gpt-6.1-sol line would be $0.023, so caching alone cuts that line by about two-thirds. Hosting, orchestration, and telephony are excluded because they apply to every architecture.

Token-billed S2S grows with call length

For gpt-realtime-2.1, assume 600 audio tokens per minute of user audio and 1,200 per minute of model audio. That rate matches the per-hour audio prices Artificial Analysis lists for GPT-Realtime-2 ($1.15 input, $4.61 output) at those token prices. Using the same speaking shares, a call-minute generates about 270 new input audio tokens and 420 output audio tokens.

Output costs 420 x $64/M = $0.027. New input costs 270 x $32/M = $0.009. Re-reading the cached context, about 11,600 tokens per minute at minute two, adds about $0.005 at $0.40/M. The total is about $0.040 per minute, and it rises slowly as the context grows. If that context misses the cache, the re-read alone would cost roughly $0.16 per minute instead of $0.005.

Cost per resolved call is the number that matters

Per-minute cost is a distraction if resolution rates differ. Use this formula:

Expected cost per call = C + (1 - r) x H

Here C is the AI cost per call, r is the resolution rate, and H is the cost of the human handling a failed call needs. Assume H = $3.50 for illustration.

For a cascade at C = $0.10 (4 minutes at $0.025) and r = 85%, the expected cost is $0.10 + 0.15 x $3.50 = $0.625. For GPT-Live at C = $0.21 (4 minutes at $0.05, plus an assumed $0.01 of backend tokens) and r = 88%, it is $0.21 + 0.12 x $3.50 = $0.63. They are nearly identical.

The break-even resolution gap is (C_b - C_a) / H. GPT-Live must beat the cascade by (0.21 - 0.10) / 3.50 = 3.1 points to pay for itself. Grok Voice at $0.32 per call needs 6.3 points. Our cost per resolution guide extends this to containment and repeat calls.

That 3-point gap is the uncomfortable part, because detecting it takes a lot of calls. More on that in the bake-off protocol below.

Hybrid patterns that work in production

The most useful architecture decision is often "which hop gets text back." Here are five hybrids, with what each costs you.

Four voice agent architectures compared: full cascade, half-cascade, speech-to-speech with a text backend, and pure speech-to-speech, showing which hops carry text

1. Half-cascade: audio in, text out

The realtime model listens, hears prosody, and returns text. Your TTS speaks it. You get exact scripted speech, brand voice choice, and stable output. LiveKit notes a quirk this avoids: some realtime models start answering in text only after loading long conversation histories.

The catch is that GPT-Live cannot do this. The LiveKit plugin docs state it "has no text-only response modality." You need a realtime model with a text modality, such as the OpenAI Realtime API.

2. S2S front, text brain behind

GPT-Live's default mode delegates reasoning and tools to a backend Responses model. Its `client` delegation mode goes further: no backend model runs, and your own code answers each delegation with `append_commentary`. That lets you keep your existing LLM, retrieval stack, and policy engine while the voice model handles turn-taking.

Know the limits from the LiveKit docs. Startup history caps at 128 messages and 8,192 tokens. Each append caps at 500 tokens. Context is append-only, so you cannot correct a wrong item. These limits are documented; they are not edge cases.

3. Cascade with S2S-grade turn-taking

Fused STT with end-of-turn detection closes much of the first-response gap while keeping text at every hop. Deepgram reports that Flux "can cut agent response latency by 200-600 ms compared to pipeline approaches." Pair it with eager end-of-turn and speculative LLM calls, and price in the extra LLM calls.

4. Route by turn type

Use S2S for open conversation and a cascade for tool-heavy or regulated steps. LiveKit's pipeline write-up mentions teams exploring this split. Be careful with the hand-off. A GPT-Live agent handoff with different instructions starts a new session, which re-sends history within the 128-message, 8,192-token startup cap. OpenAI's voice cost guide also notes that creating a WebRTC session bills 15 seconds up front, credited once the session runs, so sessions created early or abandoned still cost money.

5. Sidecar STT for the record

Run a separate streaming STT alongside the S2S model to get live captions, compliance transcripts, and turn-detection signals. LiveKit's docs recommend exactly this when you need realtime transcription or its turn detector with a realtime model. You pay for STT twice, at about $0.005 to $0.008 per minute on Deepgram's list prices.

The three configurations in code

The following is simplified. The cascade strings and `TurnHandlingOptions` follow LiveKit's published examples. The half-cascade and GPT-Live forms follow the LiveKit realtime and GPT-Live plugin docs. Check names against your installed version before use.

# Simplified, illustrative. LiveKit Agents (Python).
from livekit.agents import AgentSession, TurnHandlingOptions
from livekit.plugins import openai, silero

# 1) Cascade: text at every hop, tunable endpointing
cascade = AgentSession(
    vad=silero.VAD.load(),
    stt="deepgram/nova-3:multi",
    llm="openai/gpt-4.1-mini",
    tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
    turn_handling=TurnHandlingOptions(
        endpointing={"mode": "fixed", "min_delay": 0.3, "max_delay": 2.5},
    ),
)

# 2) Half-cascade: realtime listening, your TTS speaks
half_cascade = AgentSession(
    llm=openai.realtime.RealtimeModel(modalities=["text"]),
    tts="cartesia/sonic-3:9626c31c-bec5-4cca-baa8-f8ba9e84c8bc",
)

# 3) S2S with a backend brain: GPT-Live delegates reasoning and tools
s2s_backend = AgentSession(
    llm=openai.realtime.GPTLiveModel(
        voice="marin",
        responses_options={
            "model": "gpt-5.6-luna",
            "instructions": "Use tools when current information is required.",
        },
    ),
    vad=silero.VAD.load(),  # needed for playback cut-off on barge-in
)

That last `vad` line is a gotcha in its own right. The GPT-Live plugin docs say the session drops its default VAD for this model, and without one, "the agent plays until the model stops on its own."

Decision matrix by use case

Score the constraint that binds first. This matrix is a starting point, built from the capabilities and benchmark findings above. Validate it on your own calls.

Use caseLean towardWhyMust test before launch
Payments, collections, PCI scopeCascade, or route the payment step to a cascadeRedaction before the LLM, exact disclosuresDisclosure verbatim rate, redaction misses
Healthcare intake and schedulingCascade or half-cascadeExact wording, live transcript, audit trailEntity accuracy on names, dates, medications
Tier-1 support with many toolsS2S + backend, or cascade with a strong LLMtau-Voice shows tool tasks drive the gapDatabase end-state success under noise
Outbound qualificationCascade with fused turn detectionCost at volume, scripted openersAnswering machine handling, opener timing
Coaching, companionship, sales role-playPure S2S or S2S + backendProsody, overlap, backchannelsInterruption and backchannel handling
Multilingual, accented caller baseTest both; no defaultAccent robustness varied by provider in tau-VoiceTask success per accent cohort
Noisy environments (drive-thru, field)Cascade with noise suppressionYou can tune VAD and STT per environmentSuccess at 10 dB and 5 dB SNR
Web or app voice with good audioS2S or half-cascadeClean wideband audio, no PSTN legLatency p90, overlap quality

Use a weighted score when the call is close. Assign weights to five criteria: task success, p90 latency, overlap handling, control requirements, and cost per resolved call. Weights should sum to 100. A regulated contact center might use 35, 15, 10, 25, and 15. A consumer companion app might use 20, 20, 35, 5, and 20.

How to run a fair cascade vs speech-to-speech bake-off

Most architecture comparisons are unfair by accident. They run the S2S model over WebRTC and the cascade over the phone, or they compare a vendor dashboard's latency against a stopwatch. Use this protocol to compare like with like. Our guide to comparing voice agents on the same test cases covers the scenario design in more depth.

1. Freeze a scenario set. Write 40 scenarios that cover your top intents, two policy edge cases per intent, and at least five multi-tool tasks. Give each scenario a single correct end state, as tau-Voice does, so success is checkable without a judge.

2. Cross scenarios with conditions. Run each scenario under five conditions: clean audio, 10 dB SNR babble noise, 5 dB SNR street noise, two accent cohorts drawn from your caller base, and a scripted interruption at a fixed point. That gives 200 scenario-condition cells.

3. Use the same transport leg for both. Route both architectures through the same SIP trunk at 8 kHz G.711. A 16 kHz WebRTC path flatters whichever system gets it. See our SIP vs WebRTC guide for why.

4. Use the same tools and the same mock backend. Point both agents at identical tool schemas and a seeded test database. Score success by comparing the final database state to the expected one.

5. Measure at the caller side. Record both channels on the caller's leg. Compute response latency as the gap from caller speech end to agent speech start, using one VAD for both systems. Report p50 and p90, not means.

6. Score overlap separately. Measure agent interruption rate, yield latency after a barge-in, and selectivity on backchannels such as "mm-hmm." These are the metrics where S2S should win, so give it the chance.

7. Size the sample before you start. To detect 80% vs 85% task success at 80% power and alpha 0.05, you need about 903 calls per arm. To detect 85% vs 88%, the 3-point break-even from the cost section, you need about 2,033 per arm.

8. Run three trials per cell. Artificial Analysis averages three trials where available on its tau-Voice charts. Single runs hide variance from model sampling and simulator randomness.

9. Report cost per resolved call with confidence intervals. Use the formula above, with your real human handling cost. If the intervals overlap, the architectures tie on your data, and control requirements should decide.

Here is the sample-size and latency arithmetic in Python, so you can rerun it with your own rates:

# Two-proportion sample size and caller-side latency percentiles (illustrative)
from math import ceil
from statistics import quantiles

Z_ALPHA, Z_BETA = 1.96, 0.8416  # two-sided alpha 0.05, power 0.80

def calls_per_arm(p1: float, p2: float) -> int:
    var = p1 * (1 - p1) + p2 * (1 - p2)
    return ceil((Z_ALPHA + Z_BETA) ** 2 * var / (p1 - p2) ** 2)

print(calls_per_arm(0.80, 0.85))  # 903 (about 902.6)
print(calls_per_arm(0.85, 0.88))  # 2033

def latency_report(caller_end_ms: list[float], agent_start_ms: list[float]) -> dict:
    gaps = [a - c for c, a in zip(caller_end_ms, agent_start_ms)]
    pct = quantiles(gaps, n=100)
    return {"p50_ms": pct[49], "p90_ms": pct[89], "turns": len(gaps)}

Those sample sizes explain why so many architecture decisions rest on anecdotes. A 50-call pilot cannot separate two systems that differ by 3 to 5 points. This is where an independent evaluator helps: Evalgent runs the same scenario matrix against both stacks on the same transport leg and reports the intervals, so the decision rests on your calls rather than a leaderboard. For the statistics behind it, see our A/B testing guide for voice agents.

Non-obvious gotchas from production

These are the failure modes that show up after launch, not in demos.

  • S2S agents confirm things the caller never heard. Without truncation, the model's context holds the full interrupted turn. Test it by interrupting mid-confirmation and checking what the agent claims it said.
  • Late transcripts break transcript-keyed logic. Escalation rules, sentiment triggers, and keyword guards that read the user transcript fire after the model has already answered. Move those checks to tool calls or a sidecar STT.
  • Scripted speech is not available on full-duplex models. LiveKit's docs say GPT-Live "can't speak a script word for word." Even with a TTS attached for `session.say()`, the model can talk over it.
  • History seeding has hard caps. GPT-Live accepts 128 messages and 8,192 tokens at startup, oldest dropped first. Long callbacks with prior context lose the oldest facts silently.
  • Eager turn detection inflates LLM spend. Deepgram puts it at 50% to 70% more LLM calls. Price it into the cascade's cost line, not just the latency line.
  • Default cascades on leaderboards use older LLMs. The arena's cascaded entries run GPT-4o Mini and Gemini 2.5 Flash. A cascade with a current model is a different system.

For a broader test plan for the S2S side, see our guides to testing speech-to-speech voice agents and evaluating a speech-to-speech model. For framework-level latency differences, see LiveKit vs Pipecat latency, and for xAI's stack specifically, our Grok voice agent guide.

Frequently asked questions

Is speech-to-speech faster than a cascaded voice pipeline?

Usually, but by less than vendors claim. Independent measurements put leading S2S models at about 0.7 to 1.25 seconds model-side. A default cascade with a 500 ms endpointing delay lands near 1.8 seconds on a phone call. A tuned cascade with fused turn detection lands near 1 second. S2S's clearer advantage is handling overlap, interruptions, and backchannels.

Do speech-to-speech models reason worse than text models?

On short spoken puzzles, not anymore. Top S2S models now score 96% to 99% on Big Bench Audio, up from 66% for GPT-4o Realtime in 2024. On multi-turn tool tasks, the gap remains. In tau-Voice, voice agents kept only 30% to 45% of a text model's task success under realistic audio.

What is a half-cascade voice agent?

A half-cascade pairs a realtime audio model, which listens and returns text, with a separate TTS that speaks. You keep prosody-aware listening and gain exact scripted speech and voice choice. It requires a realtime model with a text-only response modality. GPT-Live does not offer one, according to LiveKit's plugin docs.

How much does a speech-to-speech voice agent cost per minute?

GPT-Live lists $0.05 per minute of voice time, billed per second, plus backend tokens. Grok Voice lists $0.08 per minute, plus $0.01 on a free xAI number. Token-billed gpt-realtime-2.1 works out near $0.04 per minute at typical speaking shares. A self-assembled cascade can run $0.018 to $0.029 per minute.

Which is better for regulated industries, cascading or speech-to-speech?

Cascading, in most cases. It lets you redact PII before the LLM sees it, speak exact disclosures, and keep a text trail at every hop. In any S2S design, raw audio reaches the model provider and scripted speech is limited. Many teams use S2S for open conversation and route regulated steps to a cascade.

Can I use GPT-Live with my own LLM?

Yes, through client delegation. In that mode no backend model runs. Your code receives each delegation event, works out the answer with your own LLM or database, and returns it with `append_commentary`. Function tools are ignored in this mode, so your code owns all reasoning and tool execution.

How many test calls do I need to compare two architectures?

More than most pilots run. Detecting 80% vs 85% task success at 80% power needs about 900 calls per arm. Detecting a 3-point gap, such as 85% vs 88%, needs about 2,000 per arm. Smaller tests can catch large failures but cannot rank systems that differ by a few points.

Should I start with cascading or speech-to-speech in 2026?

Start with a cascade if you need exact wording, PII redaction, or per-layer debugging. Start with S2S, ideally with a text backend, if overlap, prosody, and natural barge-in drive your outcomes. In both cases, keep the scenario suite architecture-neutral. Then you can switch later and measure the difference instead of guessing.

The bottom line

Speech-to-speech models have closed the reasoning gap on short questions and win on overlap and prosody, but tuned cascades now match them within a few hundred milliseconds on first response, cost less when unbundled, and still lead on control and multi-step tool work. Pick by the constraint that binds first, then prove the choice with a same-scenario, same-transport bake-off sized to detect the resolution gap that would pay for the more expensive option.

Related Articles