Evalgent
Back to Blog
Voice AI Testing

When Your STT, LLM or TTS Provider Goes Down: Voice Agent Fallback and Failover That Holds Up on Live Calls

Deepesh Jayal
23 min read
When Your STT, LLM or TTS Provider Goes Down: Voice Agent Fallback and Failover That Holds Up on Live Calls
On this page
4 h 22 m
OpenAI API degraded or down on Dec 11, 2024 (OpenAI write-up)
5.0 s
Default per-attempt timeout in LiveKit llm.FallbackAdapter (source)
32.1 s
Worst-case TTS hang before LiveKit tries the backup, from source defaults
700-800 ms
Silence where listeners rate a reply as less willing (Roberts & Francis, 2013)

Your voice agent depends on four companies at once: a carrier, an STT provider, an LLM provider and a TTS provider. Each has incidents every month. A chatbot that waits three seconds looks slow. A voice agent that waits three seconds sounds broken, and the caller hangs up.

This guide is for teams running their own agent on LiveKit, Pipecat or the Deepgram Voice Agent API. We read the fallback source code, not only the docs, because the defaults decide what callers hear during an outage.

What callers hear when each provider fails

Every layer fails differently, and the caller experience is the only symptom that matters. Map what the caller hears to what your logs show, because the two rarely look alike.

Failure symptom map for a cascaded voice agent showing what callers hear, what logs show, and the detection signal when telephony, STT, LLM or TTS providers fail
Layer that failsWhat the caller hearsWhat your logs showFast detection signal
Telephony / SIPRinging that never connects, a carrier error message, or a fast busyNothing in the agent logs, because the call never reached youCarrier webhook error callbacks, SIP 5xx rate, inbound call count dropping to zero
STTThe agent greets, the caller answers, then silence. The agent never replies because it never "heard" anythingVAD shows user speech, but no final transcript arrivesVAD speech with no transcript within 1-2 s
LLMDead air after the caller finishes. Sometimes half a sentence, then silenceTurn ends, LLM request open, no first tokenTime to first token over budget
TTSThe agent "thinks" and never speaks. Transcripts show agent text that was never voicedLLM text generated, zero audio bytes pushedAgent turn with text but 0 ms of audio
Your agent serversCalls connect to silence, or the carrier plays its own errorJob dispatch failures, crash loopsDispatch latency, active sessions dropping

Two patterns stand out. STT and TTS failures produce one-sided silence: a transcript-only review of a TTS outage shows a polite conversation nobody heard, the trap described in voice agents that go silent mid-call. And telephony failures leave no trace in your agent at all.

Outages are routine: what the status pages show

Provider failures are on public status pages, and the shape of each incident tells you which detection method would have caught it.

The largest recent one is OpenAI's December 11, 2024 incident. Per OpenAI's write-up, a telemetry deployment overloaded the Kubernetes control plane and broke DNS-based service discovery. The API was degraded or unavailable from 3:16 PM to 7:38 PM PST, with substantial recovery at 5:36 PM. Affected components included Chat Completions, Audio and Realtime. A single-LLM voice agent was down for over two hours.

Smaller incidents are more frequent. The Deepgram status history for 2026 includes:

  • June 15-16: some Flux sessions "established successfully but then terminated a few seconds into the session," so voice-agent flows "would have seen sessions open with no transcript returned" (incident).
  • August 25: EU flux-general-multi requests failed "on the first audio" after the WebSocket connected, causing "reconnect loops" (incident).
  • September 25: Flux TTS errors on the global endpoint while "regional endpoints (EU, India, Australia) are unaffected" (incident).
  • June 17: a Google API key change broke Gemini-backed Voice Agent requests, and Deepgram advised customers to "define multiple LLM providers" (incident).

ElevenLabs' status page shows a nine-minute US error spike across TTS, STT and telephony on July 8, 2026, and an August 3 incident that ended when "the provider resolved the underlying issue." Your "independent" backup may share an upstream with your primary.

The Gray Failure paper (Huang et al., HotOS 2017) names the pattern: differential observability, where an application sees a system as unhealthy while the system's own detector sees it as healthy. A WebSocket that connects and returns no transcript raises no error, so an adapter that triggers on errors never switches. Detect what the caller experiences, not what the provider returns.

How LiveKit's FallbackAdapter decides to fail over

LiveKit offers two mechanisms, described on its fallback strategies page. The Inference Fallback Adapter runs server-side inside LiveKit Inference and supports STT and TTS through a `fallback` list on `inference.STT` and `inference.TTS`. The Agent Fallback Adapter runs in your agent process and supports STT, LLM and TTS: `stt.FallbackAdapter`, `llm.FallbackAdapter` and `tts.FallbackAdapter`. The docs say both trigger on "any error from the primary provider, including connection failures, timeouts, HTTP errors (4xx, 5xx), and mid-stream disconnects," mark the failed provider unhealthy, and probe it in the background.

The docs stop there. The constructor defaults, which decide how long your caller waits, are in the source. Here is what the current `main` branch of livekit/agents sets.

AdapterConstructor defaultsWhat the timeout actually bounds
`llm.FallbackAdapter``attempt_timeout=5.0`, `max_retry_per_llm=0`, `retry_interval=0.5`, `retry_on_chunk_sent=False`Passed as `httpx.Timeout(5.0)`, which is a per-phase limit (connect, read gap, write, pool), not a total deadline
`stt.FallbackAdapter``attempt_timeout=10.0`, `max_retry_per_stt=1`, `retry_interval=5`Connection setup and request time; it does not watch for a stream that stays open but returns nothing
`tts.FallbackAdapter``max_retry_per_tts=2`, plus the session's TTS connect options (`timeout=10.0`, `retry_interval=2.0` by default)Each attempt; retries run on the same provider before the backup is tried

The LLM timeout is not a time-to-first-token deadline

In the LiveKit LLM stream, the per-attempt timeout becomes `httpx.Timeout(self._conn_options.timeout)`. An httpx timeout with one value applies separately to connect, read, write and pool. The read timeout is the longest gap between received bytes, not the total wait. So with the default of 5.0 s, a provider that takes 4.9 s to send its first token passes. A provider that trickles a token every 4 s never trips it. Only a hard error or a gap longer than 5 s triggers failover.

For a voice agent, 5 s of dead air after the caller stops talking is a failed turn. Cut `attempt_timeout` to just above your measured p99 time to first token, and measure that number on your own prompts, since TTFT depends on prompt length and model. If you use a reasoning model with long silent thinking, the read-gap semantics mean a tight timeout will also cut off healthy slow responses. Test both.

TTS retries the same broken provider before trying the backup

The TTS adapter's default `max_retry_per_tts=2` means three attempts on the primary before it moves on. LiveKit's base retry logic waits 0.1 s before the first retry and `retry_interval` (2.0 s) after that, and each attempt is bounded by the session's TTS timeout of 10.0 s. If the primary hangs instead of erroring, the worst case before the backup gets the text is 10 + 0.1 + 10 + 2.0 + 10 = 32.1 s. If the primary fails fast with a 503, the cost is closer to 2.1 s of retry waits plus request time. Either way, the caller hears nothing. Set `max_retry_per_tts=0` and let the backup do the retrying.

Health state lives on the adapter instance

Each adapter keeps a list of `available` flags, set to `True` in `__init__`. When a provider fails, the adapter flips it to `False`, emits `llm_availability_changed` (or the STT/TTS equivalent), and starts a background recovery task. Most LiveKit agents construct `AgentSession`, and the adapters with it, inside the job entrypoint. That means every new call starts with the primary marked healthy and pays the detection cost again on its first turn. During a 40-minute LLM outage, every call that arrives pays one full timeout before it reaches the backup.

The fix is to keep failover memory outside the call: a process-level or Redis-backed health flag that your entrypoint reads to order the provider list, so new calls start on the healthy provider.

Recovery probes use real traffic

The LLM recovery task replays the current turn's chat context against the failed provider. The TTS adapter re-synthesizes the text the caller is hearing. The STT adapter forwards the caller's live audio to a second stream on the failed provider and marks it recovered only after one non-empty final transcript. So probes carry caller content, may be billable if the provider is partly up, and a recovered primary takes over mid-call, which for TTS means the voice can switch twice in one conversation.

Partial output guards

STT switches on any error. TTS will not switch once audio has been pushed for the current segment, so the caller hears part of a sentence, then nothing. The LLM raises instead of restarting if text or tool calls were already streamed, unless you set `retry_on_chunk_sent=True`, in which case the backup regenerates the response and the caller may hear the start twice.

Pipecat: ServiceSwitcher, is_usable, and what it does not catch

Pipecat's equivalent is the `ServiceSwitcher`, a parallel pipeline where filters let frames reach only the active service. `LLMSwitcher` is the LLM-specific version, with `register_function` to register a tool handler on every LLM at once. Pass `strategy_type=ServiceSwitcherStrategyFailover` to get automatic switching. The default is `ServiceSwitcherStrategyManual`, which never switches on its own.

The behavior depends on Pipecat's error model. Every processor has `is_usable`, and errors carry an `ErrorCategory`. `AUTHENTICATION`, `AUTHORIZATION` and `INVALID_REQUEST` are permanent and make a processor unusable. `RATE_LIMIT`, `QUOTA`, `CONNECTIVITY`, `SERVER` and `UNKNOWN` are not. HTTP 401, 403, 429 and 5xx map automatically. The docs say the failover strategy switches when a service is "unable to work at all — a rejected API key, an unknown model or voice," or after "enough consecutive failures that it has stopped trying."

Three gaps follow:

  • A 503 burst may not switch you over. Under the documented semantics, transient errors keep you on the failing provider. On the `main` branch we read, `ServiceSwitcherStrategyFailover.handle_error` switched on any non-fatal error. Docs and source drift between releases, so pin your version and test it.
  • Slow is not an error. A provider that answers in 6 s emits no `ErrorFrame`. You need your own watchdog.
  • No automatic recovery. The failed service stays out until you call `set_usable(True)` and switch back.

If you want transient errors to trigger failover, subclass the strategy. The docs describe `handle_error(error)` as the override point and `error.processor.is_usable` as the signal. This is an illustrative sketch:

# Illustrative: fail over on transient provider errors, not only permanent ones.
from pipecat.pipeline.service_switcher import ServiceSwitcher, ServiceSwitcherStrategyFailover
from pipecat.utils.errors import ErrorCategory

TRANSIENT = {
    ErrorCategory.SERVER,
    ErrorCategory.CONNECTIVITY,
    ErrorCategory.RATE_LIMIT,
    ErrorCategory.QUOTA,
}

class FailoverOnTransient(ServiceSwitcherStrategyFailover):
    async def handle_error(self, error):
        finished = error.processor is not None and not error.processor.is_usable
        if finished or getattr(error, "category", None) in TRANSIENT:
            return await super().handle_error(error)  # switch to next service
        return None  # APPLICATION errors (your tool code) never trigger a switch

stt_switcher = ServiceSwitcher(
    services=[deepgram_stt, backup_stt],
    strategy_type=FailoverOnTransient,
)

@stt_switcher.strategy.event_handler("on_service_switched")
async def on_switched(strategy, service):
    metrics.increment("stt_failover", tags={"to": service.name})

Also set the pipeline policy deliberately. `PipelineWorker(pipeline, processor_unusable_policy=ProcessorUnusablePolicy.CONTINUE)` is the default and suits a switcher. `END` is what Pipecat's examples use, and it hangs up on the caller the moment one service is unusable. More on Pipecat test setup is in our Pipecat voice agent testing guide.

Deepgram Voice Agent API: per-turn fallback with no memory

If you use Deepgram's managed Voice Agent API, the `think` field accepts an array of providers. Per Deepgram's LLM models docs, the agent sends each request to the first provider, and on error or timeout it emits a `THINK_REQUEST_FAILED` warning over the WebSocket and tries the next. If all fail, you get `FAILED_TO_THINK` and "the turn produces no LLM response." The docs also say: "The fallback is per-request — each new conversational turn starts again from the first provider."

There is no unhealthy flag, so during a primary outage every turn pays the primary's failure time. If the primary hangs until a timeout, every turn has dead air. Mixing provider types (an `open_ai` primary with an `anthropic` fallback) is allowed and gives you independent infrastructure. Listen for `THINK_REQUEST_FAILED`, and if it repeats, send an Update Think message to reorder providers for the rest of the call.

Timeouts are the real trigger: build a dead-air budget

Errors are the easy case. A slow or silent provider is decided entirely by your timeout values, so set them from what humans tolerate, not from HTTP library defaults.

Two research results give you the budget. Stivers and colleagues studied turn-taking across 10 languages (PNAS, 2009) and found that every language's distribution of response gaps peaks between 0 and 200 ms, with average offsets within 500 ms. Roberts and Francis (JASA, 2013) played 380 listeners dialogues where an identical "yes" followed a request after silences from 200 to 1,200 ms. Ratings of the responder's willingness dropped noticeably at 600 ms, with a statistically significant difference between 700 and 800 ms. In production terms, the caller starts reading meaning into silence after roughly 0.6-0.8 s. Your failover has to finish, or be masked by a filler, before then.

Dead-air budget timeline comparing LiveKit default fallback timeouts with tuned timeouts and a filler phrase, showing where caller tolerance ends at 600 to 800 milliseconds

The budget for a failed-then-recovered turn is:

`dead air = endpointing delay + failure detection time + backup TTFT + TTS time to first byte`

Plug in illustrative numbers. Assume 400 ms endpointing, 500 ms backup TTFT and 150 ms TTS first byte.

ConfigurationDetection timeDead air on a hung primaryCaller experience
LiveKit LLM default5,000 ms400 + 5,000 + 500 + 150 = 6,050 msCaller says "hello?" or hangs up
Tuned `attempt_timeout=1.5`1,500 ms400 + 1,500 + 500 + 150 = 2,550 msNoticeable pause
Tuned plus filler at 700 ms1,500 msSilence ends at 700 ms; answer at 2,550 msSounds like the agent is checking
Hedged request at p95 TTFTp95, for example 900 msAbout 400 + 900 + 500 + 150 = 1,950 msClose to normal

The last row comes from Dean and Barroso's The Tail at Scale (CACM, 2013). A hedged request sends a second copy once the first has been outstanding past the 95th-percentile expected latency, limiting extra load to about 5%. In their BigTable benchmark, hedging after 10 ms cut 99.9th-percentile latency from 1,800 ms to 74 ms with 2% more requests. It only helps when the slowdown does not hit both replicas at once. For voice, hedge the LLM call to a second provider when TTFT passes your p95 and cancel the loser. None of these frameworks hedge for you; build it in a custom LLM node.

Here is an illustrative tuned LiveKit configuration. Check the values against your own TTFT and TTS distributions, collected as in our LiveKit latency debugging guide.

# Illustrative LiveKit Agents config with tightened failover timing.
from livekit.agents import AgentSession, APIConnectOptions, llm, stt, tts
from livekit.agents.voice.agent_session import SessionConnectOptions
from livekit.plugins import assemblyai, cartesia, deepgram, elevenlabs, openai, silero

llm_fb = llm.FallbackAdapter(
    [
        openai.LLM(model="gpt-4.1-mini"),
        openai.LLM.with_azure(model="gpt-4.1-mini"),  # same model, separate infrastructure
    ],
    attempt_timeout=1.5,     # just above measured p99 TTFT; default is 5.0
    max_retry_per_llm=0,     # fail over instead of retrying the broken provider
    retry_on_chunk_sent=False,
)

stt_fb = stt.FallbackAdapter(
    [deepgram.STT(model="nova-3"), assemblyai.STT()],
    attempt_timeout=3.0,     # default is 10.0
    max_retry_per_stt=0,     # default is 1, with a 5 s retry interval
)

tts_fb = tts.FallbackAdapter(
    [cartesia.TTS(voice=PRIMARY_VOICE_ID), elevenlabs.TTS(voice_id=MATCHED_VOICE_ID)],
    max_retry_per_tts=0,     # default is 2 retries on the same provider
)

session = AgentSession(
    stt=stt_fb,
    llm=llm_fb,
    tts=tts_fb,
    vad=silero.VAD.load(),
    transcription_timeout=2.0,  # emit user_transcription_timeout when STT goes quiet
    conn_options=SessionConnectOptions(
        tts_conn_options=APIConnectOptions(max_retry=0, timeout=2.0),
    ),
)

for adapter, kind in ((llm_fb, "llm"), (stt_fb, "stt"), (tts_fb, "tts")):
    adapter.on(
        f"{kind}_availability_changed",
        lambda ev, kind=kind: metrics.gauge(f"{kind}_provider_available", int(ev.available)),
    )

The `transcription_timeout` line closes the gray-failure gap for STT. Per LiveKit's events reference, the session then emits `user_transcription_timeout` when VAD detected speech but STT returned no final transcript in time. That is the symptom of the June 2026 Flux incident, which no error-based adapter would catch. Ask the caller to repeat, and if it fires twice in a row, restart STT on the backup.

Mid-turn versus next-turn failover

Where in the turn the switch happens decides what the caller hears.

Failover pointExampleCaller hearsBest for
Before first byteLLM connect fails, TTS returns 503 before audioA longer pause, then a normal answerEvery layer; this is where most failover should happen
Mid-stream, restart`retry_on_chunk_sent=True` after the LLM dies mid-sentenceStart of the sentence repeated by the backupShort answers where a repeat is acceptable
Mid-stream, abandonTTS dies after audio started; LiveKit plays the partial and stopsHalf a sentence, then silenceNothing; follow it with a recovery line
Next turnStream errors, agent waits for the callerSilence until the caller speaks againFallback of last resort

The abandon case needs a recovery step, because the chat context may record the full text while the caller heard half. Speak a short bridge on the backup voice: "Sorry, you cut out on my end. I said your appointment is Tuesday at 3."

For tool calls, LiveKit's default refuses to restart on another model once a tool call has streamed, which protects you from running a payment twice. Keep that default for tools with side effects, and make handlers idempotent with a call-scoped request ID. Our tool calling test cases cover the checks.

Parity: voice, keyterms and tool schemas

A backup that is up but behaves differently is a different failure, and one parity problem is hidden in the adapter code.

The adapter advertises only what every provider supports

LiveKit's STT adapter sets capabilities as the intersection of its providers: `interim_results=all(...)`, `diarization=all(...)`, and aligned transcripts only if all support them. The TTS adapter does the same for aligned transcripts. Add a backup without interim results, and the adapter reports none even while the healthy primary serves 100% of traffic, which can change interruption behavior the day you add the backup. Compare the adapter's `capabilities` against the primary alone. The TTS adapter also resamples every provider to the highest sample rate among them.

Keyterms and language

The STT adapter forwards session keyterms to every provider, and "unsupported ones warn-and-skip internally." A backup without keyterm support loses the boost on product names, drug names and SKUs during failover, so entity accuracy drops while word error rate looks fine. Test the backup with the method in our STT entity accuracy guide.

Voice mismatch on TTS fallback

A caller who hears a new voice mid-call may think they were transferred. Rank your options:

1. Same provider, different region. In Deepgram's September 25 incident, regional endpoints were unaffected while the global one failed. A regional endpoint keeps the exact voice. Check data residency first.

2. Same cloned voice on two providers. LiveKit notes that with LiveKit Inference custom voices, "each cloned voice is cloned to more than one provider," so fallback across providers is automatic.

3. A matched stock voice. Pick the closest gender, age and pace on the backup, then run a listening test.

4. Sticky failover. Once a call switches voices, keep it on the backup for the rest of the call, even if the primary recovers. LiveKit's recovery probe will otherwise restore the primary mid-call and switch the voice a second time.

LLM prompts and tool schemas

The same prompt behaves differently on another model family, and tool schema handling differs: LiveKit's OpenAI plugin uses strict tool schemas by default, while several OpenAI-compatible helpers turn strict mode off. Run your full regression suite against the backup model. A same-model backup on separate infrastructure, such as OpenAI direct plus Azure OpenAI, avoids most drift. Our guide to swapping voice agent providers has a parity checklist.

Graceful degradation when every provider fails

All providers in a chain can fail at once: a shared cloud region, DNS, or your own egress. Plan the last line per layer so no caller sits in silence.

What is downDegradation moveHow
TTS (all)Play pre-recorded audioLiveKit documents `session.say(text, audio=audio_frames_from_file(path))`, which bypasses TTS
STT (all)Switch to keypad inputDTMF tones travel as telephony events, not speech, so "press 1 for a callback" still works
LLM (all)Scripted pathPre-recorded "I can't look that up right now. I'll text you a link, or press 0 for a person."
Agent serversCarrier-level fallbackTwilio Elastic SIP Trunking's disaster recovery URL returns TwiML when no origination URI answers
EverythingHuman transferWarm transfer to a queue with the reason attached; see our call transfer guide

Write the scripts before the outage: say what happened plainly, offer one next step, and never promise a time.

  • Filler at 700 ms of silence: "One moment while I pull that up."
  • Degraded mode: "I'm having trouble with my system right now. I can text you a link to finish this, or connect you to someone. Which would you like?"
  • Callback offer: "I can call you back within the hour at this number. Press 1 for yes."

On LiveKit, the events reference shows the pattern: when `ev.error.recoverable` is `False`, play the pre-recorded file. For LLM and TTS errors you can set `ev.error.recoverable = True` to keep the session alive, since those components are recreated per response. STT errors need `session.update_agent(session.current_agent)` to restart the stream. Without a handler, an unrecoverable error closes the session and the caller hears a dead line.

At the telephony layer, Twilio's voice failover best practices list multiple origination SIP URIs with priority and weight, a disaster recovery URL that returns TwiML, and fallback webhook URLs in another region or cloud. Twilio's October 16, 2026 IE1 maintenance notice warns of interruptions of up to 90 seconds and says customers "configured for failover to DE1 or alternate regions should experience no service disruption" (notice).

The availability math: chain versus fallback

A cascaded voice agent is a series system. Every hop must work for the call to work, so availabilities multiply. A fallback pair is a parallel system: it fails only if both fail.

  • Series chain: `A_chain = A_tel × A_stt × A_llm × A_tts × A_app`
  • Fallback pair, independent: `A_pair = 1 − (1 − a)(1 − b)`
  • Fallback pair, correlated: `A_pair = 1 − (1 − a) × c`, where `c` is the probability the backup is also down given the primary is down

Plug in illustrative numbers. Assume 99.95% for telephony (the figure in Twilio's API SLA for standard customers, with 99.99% on Enterprise Edition), 99.9% for each of STT, LLM and TTS, and 99.95% for your own agent servers. Assume 30,000 calls per month, a plausible mid-size volume, and that outages hit traffic in proportion to time.

SetupChain availabilityDowntime per 30-day monthCalls affected per month
No fallback99.60%About 173 minAbout 120
STT, LLM, TTS fallback, correlated (c = 0.2)99.84%About 69 minAbout 48
STT, LLM, TTS fallback, independent99.90%About 43 minAbout 30
Bar chart of voice agent availability and affected calls per month with no fallback, correlated fallback and independent fallback, showing telephony and agent servers dominate once model fallback exists

Three things fall out of the table. First, model-layer fallback cuts affected calls by roughly 60-75% in this example.

Second, once the model layers have backups, the unprotected hops dominate. Telephony and your agent servers at 99.95% each account for almost all of the remaining 30 calls (0.9995 × 0.9995 = 99.90%). Adding a fourth LLM provider before a second SIP trunk or a second server region is work on the wrong layer.

Third, correlation erases much of the gain. Deepgram's July 24, 2026 incident came from an AWS us-west-2 networking event. Two providers in the same cloud region can fail together, so ask where each one hosts.

The table understates impact in two ways. Calls already in progress when an outage starts also fail: at 30,000 calls over 22 business days of 10 hours, arrivals average about 2.3 per minute, so with 4-minute calls about 9 are in flight at any moment. And if your adapters are built per call, every call arriving during a 43-minute LLM outage pays one full timeout on its first turn, about 98 calls with a long pause. For turning these numbers into customer-facing targets, see voice agent SLA requirements.

What standby providers cost

Most STT, LLM and TTS providers bill per use, so an idle backup costs nothing until it serves traffic. With the numbers above, failover runs about 43-69 minutes a month at about 9 concurrent calls, roughly 400-620 call-minutes. Even at twice the primary's per-minute price, that is small next to an hour of failed calls. The real costs are elsewhere:

  • Rate limits. An emergency-only account often sits on a low usage tier or concurrent-stream limit, returns 429s at full load, and becomes the outage. Size it for 100% of traffic.
  • Contract minimums. Committed-spend agreements make a second provider a real line item; see running multiple voice agent vendors.
  • Testing time. Every prompt, tool and voice change must pass on two models and two voices. Skipping this is how backups end up silently broken.
  • Compliance scope. Each provider is another subprocessor receiving call audio or transcripts.

Monitoring and alerts for failover

A fallback that fires silently hides incidents. Log failover per call, as in what to log on every voice agent call, and alert on it.

SignalSourceAlert when
Provider availability changesLiveKit `*_availability_changed` events, Pipecat `on_service_switched`, Deepgram `THINK_REQUEST_FAILED`Any primary marked unavailable; page if it stays unavailable 5 min
Share of turns served by backupActive instance per turnAbove 1% for 10 min outside a known incident
First-turn TTFT p95Per-turn metricsAbove your attempt timeout minus margin
Speech with no transcriptLiveKit `user_transcription_timeout`More than 2% of turns in 5 min
Agent turns with text and zero audioTTS metricsAny sustained rate
Max silence per callAudio timelinep95 above 2.5 s
Calls closed by errorSession `close` events with reason `error`Above baseline
Inbound call volumeCarrierDrops below expected for time of day

For paging, use multi-window burn-rate alerts from Google's SRE workbook: with a 30-day budget, a burn rate of 14.4 for one hour consumes 2% of the month's budget, checked over a 1-hour and a 5-minute window. Define the SLO on caller experience, such as "turns with first audio within 2.5 s," because provider success rates stay green through gray failures.

Also watch for retry storms. The Metastable Failures paper (Bronson et al., HotOS 2021) describes systems that stay overloaded after the trigger is gone because retries keep load high. SDK retries inside adapter retries inside your own loop do exactly that to a recovering provider. Keep retries in one layer, add jitter, and prefer failing over.

How to test voice agent fallback before your callers do

Failover code that has never fired in a test will not work in an outage. Run these steps in a voice agent staging environment that mirrors production providers.

1. List every dependency on the call path. Carrier, SIP trunk, agent servers, STT, LLM, TTS and every tool API, each with its backup and degradation move. Any row without a backup is a known single point of failure.

2. Build a chaos endpoint for each fault shape. Real outages are not only 500s: simulate hangs, slow first tokens, trickling streams and mid-stream disconnects. This illustrative OpenAI-compatible server uses `aiohttp`.

# chaos_llm.py - OpenAI-compatible chat endpoint that misbehaves on purpose.
# Run: python chaos_llm.py ; point a client at http://127.0.0.1:8099/<mode>/v1
import asyncio, json
from aiohttp import web

def sse(content):
    chunk = {"id": "chaos", "object": "chat.completion.chunk", "created": 0, "model": "chaos",
             "choices": [{"index": 0, "delta": {"role": "assistant", "content": content},
                          "finish_reason": None}]}
    return f"data: {json.dumps(chunk)}\n\n".encode()

async def completions(request):
    mode = request.match_info["mode"]
    if mode == "http503":
        return web.json_response({"error": {"message": "overloaded"}}, status=503)
    if mode == "http429":
        return web.json_response({"error": {"message": "rate limited"}}, status=429)
    if mode == "hang":
        await asyncio.sleep(60)                     # accept the request, never answer
    resp = web.StreamResponse(headers={"Content-Type": "text/event-stream"})
    await resp.prepare(request)
    words = ["Your ", "order ", "ships ", "on ", "Tuesday."]
    for i, w in enumerate(words):
        if mode == "slow_first" and i == 0:
            await asyncio.sleep(4.0)                # first token late, under a 5 s read timeout
        if mode == "trickle":
            await asyncio.sleep(1.4)                # never exceeds a 1.5 s read-gap timeout
        if mode == "die_mid" and i == 2:
            request.transport.close()               # drop the connection mid-sentence
            return resp
        await resp.write(sse(w))
    await resp.write(b"data: [DONE]\n\n")
    return resp

app = web.Application()
app.router.add_post("/{mode}/v1/chat/completions", completions)
web.run_app(app, host="127.0.0.1", port=8099)

3. Measure time to first token through the real adapter. Drive LiveKit's `llm.FallbackAdapter` directly with the chaos endpoint as primary and a healthy endpoint as backup, and record when the first text arrives. This runs in CI in seconds. It is illustrative.

# test_llm_failover.py - pytest -q (requires pytest-asyncio)
import time, pytest
from livekit.agents import llm
from livekit.plugins import openai

def chaos(mode):
    return openai.LLM(model="chaos", base_url=f"http://127.0.0.1:8099/{mode}/v1", api_key="test")

async def first_token_s(primary_mode, attempt_timeout=1.5):
    fb = llm.FallbackAdapter([chaos(primary_mode), chaos("ok")],
                             attempt_timeout=attempt_timeout, max_retry_per_llm=0)
    ctx = llm.ChatContext.empty()
    ctx.add_message(role="user", content="Where is my order?")
    t0 = time.perf_counter()
    async with fb.chat(chat_ctx=ctx) as stream:
        async for chunk in stream:
            if chunk.delta and chunk.delta.content:
                return time.perf_counter() - t0
    return None

@pytest.mark.asyncio
@pytest.mark.parametrize("mode", ["http503", "http429", "hang"])
async def test_failover_within_budget(mode):
    t = await first_token_s(mode)
    assert t is not None and t < 2.0, f"{mode}: first token after {t}"

@pytest.mark.asyncio
async def test_trickle_is_not_detected():
    # Documents a known gap: a trickling primary never trips a read-gap timeout.
    t = await first_token_s("trickle")
    assert t is not None and t > 1.0

The trickle test is there on purpose. It shows that a provider sending one token every 1.4 s keeps the call on the primary with a 1.5 s timeout. If that matters for your use case, it is the argument for a TTFT watchdog or a hedged request.

4. Inject network faults, and use DROP, not REJECT. A rejected connection fails in milliseconds and makes failover look fast. A dropped packet hangs until a connect timeout, which is what a real network partition looks like. On a Linux staging host:

# Blackhole a provider: packets vanish, clients wait for their connect timeout.
for ip in $(dig +short api.deepgram.com); do sudo iptables -A OUTPUT -d "$ip" -j DROP; done
# Fast-fail variant for comparison: clients get connection refused immediately.
for ip in $(dig +short api.deepgram.com); do sudo iptables -A OUTPUT -d "$ip" -j REJECT; done
# DNS failure variant: point the hostname at an unroutable address.
echo "10.255.255.1 api.deepgram.com" | sudo tee -a /etc/hosts
# Clean up after each run.
sudo iptables -F OUTPUT

Provider IPs change, so re-resolve before each run, and never run this on a host that serves production calls.

5. Run a test matrix with real calls, not only unit tests. Unit tests prove the adapter switches. Only end-to-end calls prove what the caller hears. Place scripted calls through your real phone number into staging and inject faults during the call:

DimensionValues
ComponentSTT, LLM, TTS, tool API, SIP trunk
Fault503 fast, 429, DROP blackhole, slow first byte, trickle, mid-stream disconnect, bad API key, connects then silent
Call phaseGreeting, mid-conversation, during a tool call, while the agent is speaking
Repetitions20 per cell; 60 for the cells that matter most

Five components × eight faults × four phases is 160 cells. Prune the cells that cannot happen, such as a trickling SIP trunk, and you will have about 100.

6. Score each call against pass thresholds. These are the thresholds we suggest as a starting point. Tune them to your use case:

MetricDefinitionPass threshold
Max dead airLongest gap from caller end-of-turn to first agent audio2.5 s or less
Silence maskedFirst filler or answer audio after end-of-turn0.8 s or less
Fault-call completionTask completion with fault ÷ completion without fault95% or more
Voice switchesDistinct TTS voices heard in one call1 or fewer switches
Repeated speechAny sentence spoken twice0
Error closuresCalls ended by an unrecoverable error while a backup was healthy0

On sample size: if 20 of 20 calls pass a cell, the 95% upper bound on that cell's true failure rate is still about 14%. At 60 of 60, it is about 5%. Run more calls on the cells that cover your highest-volume flows.

7. Hold a game day every quarter. At low traffic, route a small share of calls to a deliberately broken primary for 10 minutes. Check that alerts fire, on-call knows the dashboard, and the degradation lines play.

8. Retest after every provider or prompt change. A new backup model, voice or tool can break parity without touching fallback code. Run the step 3 tests on every deploy and the call matrix on provider changes.

This is where independent evaluation fits for a team that tests by hand. Evalgent places scripted calls into your staging number, injects the faults in this matrix, and scores dead air, voice switches and task completion from the audio the caller hears, as a pre-launch audit or a regression check on each provider change. In production, call scoring can flag calls where failover fired so you can check the backup's real quality.

Frequently asked questions

What is voice agent fallback?

Voice agent fallback is the automatic switch from a failing STT, LLM, TTS or telephony provider to a backup during a live call. It has three parts: detecting the failure through an error or timeout, moving the request to a backup, and covering the gap so the caller does not hear silence. Detection speed decides most of the caller experience.

Does LiveKit's FallbackAdapter detect a slow LLM?

Only partly. The LLM adapter's attempt_timeout, 5.0 s by default, is applied as an httpx timeout, which limits each phase and the gap between received bytes, not total time to first token. A provider that sends its first token in 4.9 s passes. Lower attempt_timeout to just above your measured p99 TTFT, or add a watchdog.

Why does my agent stay silent when the TTS provider fails?

With LiveKit defaults, tts.FallbackAdapter retries the failing provider twice before trying the backup, each attempt bounded by a 10 s timeout. A hanging provider can hold the turn for about 32 s. Set max_retry_per_tts=0, lower the session TTS timeout, and keep pre-recorded audio ready for when every TTS provider is down.

How does Pipecat handle provider failover?

Pipecat uses ServiceSwitcher or LLMSwitcher with ServiceSwitcherStrategyFailover. Per the docs, it switches when the active service becomes unusable, such as a rejected key, an unknown model, or repeated failures. Transient 5xx errors and slow responses may not trigger it. Subclass the strategy to switch on transient categories, and handle recovery with set_usable(True).

Will callers notice when the TTS voice changes?

Often, yes. A mid-call voice change can sound like a transfer or a different person. Prefer a regional endpoint of the same provider, a voice cloned on two providers, or a closely matched stock voice. Once a call switches, keep it on the backup voice so a recovering primary does not switch it back.

Should the backup LLM be a different model family?

Use a different provider for independence, but keep the model as close as possible. The same model through a second host, such as OpenAI direct plus Azure OpenAI, avoids prompt and tool-schema drift. A different family needs its own regression run on your prompts and tools. Check that the two hosts do not share an upstream region.

How much does a standby provider cost?

Usage cost is small because most providers bill per use and the backup only runs during failover. The larger costs are higher rate limits on the backup account, contract minimums, probe traffic, compliance agreements, and testing every change on two providers. Budget engineering time for parity testing.

How do I test failover without breaking production?

Use a staging environment with production providers. Point primaries at a chaos endpoint that returns 503s, hangs, trickles or disconnects mid-stream. Blackhole provider IPs with iptables DROP rather than REJECT, place scripted phone calls during each fault, and score dead air, voice switches and task completion. Run a short game day each quarter.

The bottom line

Voice agent fallback is mostly a timing problem: the backup exists in your config, but the caller hears whatever happens during the seconds your framework's default timeouts allow. Tighten those timeouts, cover the gap with recorded audio and fillers, protect the telephony and server layers too, and inject real faults on scripted calls until the dead-air numbers meet your bar.

Related Articles