Open door for builders.
Deepgram Flux vs Silero VAD vs a Turn-Detector Model: Which End of Turn Detection Voice Agent Method to Use

On this page
Every voice agent answers the same question several times per turn: is the caller done? Get it wrong one way and the agent talks over someone reading a card number. Get it wrong the other way and every reply arrives after an awkward beat. This guide compares the five ways you can answer that question in 2026, using vendor docs, the benchmarks that exist, and the research that explains why each method fails where it does.
The five end-of-turn methods and how each one works inside
There are five families. They differ in what signal they read, and therefore in where they break.
| Family | Example | Signal it reads | Where it breaks |
|---|---|---|---|
| Silence timer | Silero VAD plus a timeout, OpenAI `server_vad` | Energy and voice probability, then a clock | Mid-turn pauses longer than the timer |
| STT-integrated | Deepgram Flux, Nova-3 `endpointing` | The recognizer's own acoustic and text state | Vendor language coverage, vendor-coupled behavior |
| Text turn model | LiveKit `MultilingualModel` (deprecated) | The transcript so far, plus chat context | Waits for the transcript; blind to intonation |
| Audio turn model | LiveKit `TurnDetector` v1, Pipecat Smart Turn v3.2 | Raw audio of the current turn | Audio it was not trained on, such as 8 kHz phone lines |
| S2S semantic VAD | OpenAI Realtime `semantic_vad` | The model's classifier on the user's audio | Little control, few knobs, opaque |
1. Silence timers: Silero VAD and friends
A voice activity detector scores short audio windows for "is this speech." Silero scores 32 ms windows at 8 kHz or 16 kHz and outputs a probability. A threshold turns that into speech or silence. A timer then declares the turn over after enough consecutive silence.
The defaults differ by framework, and they matter:
- LiveKit Silero plugin (docs): `activation_threshold=0.5`, `min_silence_duration=0.55` s, `min_speech_duration=0.05` s, `prefix_padding_duration=0.5` s. Only 8,000 and 16,000 Hz are supported.
- Pipecat `VADParams`: `confidence=0.7`, `start_secs=0.2`, `stop_secs=0.2`, `min_volume=0.6`. Pipecat lowered `stop_secs` from 0.8 to 0.2 because the turn model now makes the final call. We covered those internals in our Pipecat Smart Turn guide.
- OpenAI Realtime `server_vad` (docs): `threshold`, `prefix_padding_ms`, `silence_duration_ms`. The documented example uses 0.5, 300 ms and 500 ms.
The Silero threshold is not an end-of-turn knob. It decides what counts as speech. Raising it from 0.5 to 0.7 filters some noise but also drops soft callers. Silero scores any human voice highly, so a TV in the background still passes. The timer is the end-of-turn knob, and it is a blunt one.
The research explains why it is blunt. Heldner and Edlund (Journal of Phonetics, 2010) found that pauses inside a turn are often longer than gaps between turns. A 500 ms threshold caught roughly half of between-speaker gaps but also more than half of within-speaker pauses in their Map Task corpora. No single timer separates the two. Our post on VAD misfires covers the other half of the problem, where VAD triggers on non-speech.
2. STT-integrated end of turn: Deepgram Flux and Nova-3
Deepgram Flux puts turn detection inside the recognizer. It streams over `wss://api.deepgram.com/v2/listen` and emits `TurnInfo` messages whose `event` is one of `StartOfTurn`, `Update`, `EagerEndOfTurn`, `TurnResumed` or `EndOfTurn` (state machine docs). Each message carries an `end_of_turn_confidence` value. `Update` arrives about every 0.25 s of audio.
Three parameters control it (configuration docs):
| Parameter | Range | Default | What it does |
|---|---|---|---|
| `eot_threshold` | 0.5 to 1.0 | 0.7 | Confidence needed to fire `EndOfTurn`. 1.0 disables natural detection so you drive turns with `ForceEndTurn`. |
| `eager_eot_threshold` | 0.3 to 0.9 | not set | Enables `EagerEndOfTurn` and `TurnResumed`. Must be less than or equal to `eot_threshold`. |
| `eot_timeout_ms` | 500 to 60,000 | 5,000 | Silence after which `EndOfTurn` fires regardless of confidence. |
Two details in the state machine are useful for measurement. Every `EndOfTurn` carries a `trigger` field: `model`, `manual` or `timeout`. A rising share of `timeout` triggers means the model is unsure and callers are sitting through up to 5 seconds of dead air. And the `EndOfTurn` transcript always matches the preceding `EagerEndOfTurn` transcript. If the words change, Flux sends `TurnResumed` first. That guarantee is what makes speculative generation safe to cache.
Flux comes as `flux-general-en` and `flux-general-multi`. The multilingual model covers English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch, and accepts a `language_hint` list.
Nova-3 has no turn model. It has two silence features (endpointing, utterance end):
- `endpointing` defaults to 10 ms of silence and sets `speech_final=true`. It finalizes the transcript. It does not mean the caller is done.
- `utterance_end_ms` has a minimum of 1,000 ms and requires `interim_results=true`. Deepgram's own docs warn that it fires on a gap even when speech continues, which makes it a poor end-of-turn signal for agents.
If you are deciding between the two models for transcription quality as well as timing, see Deepgram Flux vs Nova-3 and our notes on Deepgram finalization latency.
3. Text-based turn models
A text model reads the transcript so far and predicts whether the sentence is complete. TurnGPT (Ekstedt and Skantze, Findings of EMNLP 2020) showed that a transformer language model can use dialog context and pragmatic completeness to predict turn shifts better than prior baselines.
LiveKit shipped this idea as its `MultilingualModel`, built on Qwen2.5-0.5B-Instruct. The LiveKit docs list 396 MB on disk, about 50 to 160 ms per turn on CPU, and under 500 MB of RAM. LiveKit reports a 99.3% true positive rate across languages, and true negative rates from 85.1% (Italian) to 96.3% (Hindi), with English at 87.0%. In plain terms, about 13% of English mid-turn pauses in LiveKit's test set were judged finished.
The text model is now deprecated and slated for removal in LiveKit Agents 2.0. The reason is structural. It cannot run until the transcript arrives, so STT latency lands on every turn. And identical words can be finished or unfinished depending on intonation. LiveKit's own example is "I would like to order one large pizza..." which reads the same whether or not "and a garlic bread" follows.
4. Audio-based turn models: LiveKit Turn Detector v1 and Smart Turn v3
LiveKit `TurnDetector` (docs) is now the default in `AgentSession`. It encodes the caller's audio directly, with a semantic branch (audio encoder plus adapter into an LLM) and an acoustic branch for timing and prosody, fused into one prediction (LiveKit blog). It looks only at the current user turn, so it does not need chat history.
- `v1` runs on LiveKit Inference and is free for agents deployed to LiveKit Cloud. Local development gets a monthly free allowance.
- `v1-mini` runs on your CPU, free anywhere. LiveKit recommends compute-optimized instances, not burstable ones, to avoid inference timeouts.
- It supports 14 languages: English, Arabic, German, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Dutch, Portuguese, Turkish and Chinese.
- It requires VAD with `min_silence_duration` of at least 0.25 s, or the session raises `ValueError` at start.
- `unlikely_threshold` overrides the per-language calibrated threshold, as a scalar or a per-language dict.
- If `v1` does not answer within about a second, the session commits the turn anyway and falls back to `v1-mini` for the rest of the session. That fallback is sticky.
Pipecat Smart Turn v3.2 (repo) is Whisper Tiny plus a linear classifier, about 8M parameters. It ships as an 8 MB int8 CPU build and a 32 MB fp32 GPU build. It covers 23 languages and takes up to 8 seconds of 16 kHz mono audio. The README reports about 65 ms per inference on a Pipecat Cloud 1x instance. It runs once per VAD stop, not continuously.
5. Speech-to-speech semantic VAD
The OpenAI Realtime API offers `semantic_vad` alongside `server_vad`. A classifier scores whether the user is done. When the probability is low, the model waits for a timeout; when high, it responds. The only tuning knob is `eagerness`: `low`, `medium`, `high` or `auto` (equal to `medium`). The API reference gives the maximum timeouts as 8 s, 4 s and 2 s respectively.
You can also hand turn-taking back to the client. LiveKit supports this for the OpenAI Realtime API, but not for the Gemini Live API, which keeps turn-taking on its own signals (LiveKit turns overview). Our guide to OpenAI Realtime and Gemini Live on phone calls covers the telephony side.
What the research says about end-of-turn tradeoffs
Five findings are worth carrying into production. They come from papers most voice teams have not read.
1. The latency cost of fewer cut-offs is logarithmic, and you can calculate it. Raux and Eskenazi (SIGDIAL 2008) studied 18,013 user turns from Let's Go, a public bus-information phone line in Pittsburgh. Callers used cell phones and landlines with cars, crying infants and TVs in the background. They found pause durations follow an exponential distribution, so false-alarm rate and latency follow FA = e^(β + α·L), with an R² of 0.99 on log scale. The practical meaning: every halving of your cut-off rate costs a fixed number of milliseconds, not a fixed percentage. Their 700 ms fixed timer produced about a 4% turn cut-in rate.
2. Context-dependent thresholds beat any single threshold. The same paper clustered silences by dialog features and set a separate threshold per cluster. Semantics and timing were the most informative features; prosody added little once those were included. Deployed live, median latency fell from 957 ms to 561 ms. After confirmation questions ("Leaving from the airport, is this correct?") it fell from 965 ms to 324 ms. After open questions it was slightly slower, 1,047 ms vs 993 ms, which is correct behavior: open answers contain more pauses. You can do the same today by changing `eot_threshold` or VAD settings per prompt type, which the code later in this post shows.
3. Putting the endpointer inside the recognizer halves latency. Chang et al. (ICASSP 2019, Google) folded the endpointer into an end-to-end ASR model on a voice search task. They report no quality loss and more than 2x lower latency compared with a separate endpointer. This is the design idea behind Flux. Voice search queries are short, though, and say little about long support turns.
4. On 8 kHz phone calls, text may hurt more than it helps. A September 2026 paper, "Less can be More" (arXiv 2609.11066), trained the same lightweight streaming classifier on acoustic, prosodic and text inputs. The data was real 8 kHz stereo calls across banking, insurance, retail and telecom, with a median turn of 10 to 11 seconds. Acoustic plus prosody was best: utterance F1 0.93 with 7.8% false alarms at 400 ms median latency. Adding text raised premature detections. Text alone averaged more than five premature triggers per utterance, because long turns contain many grammatically complete sentences before the speaker yields. The caveat is that these were human-to-human calls, and the text branch was a BERT encoder on streaming ASR, not an LLM.
5. Text models know when a sentence could end, not when a person will stop. TurnGPT showed language models capture syntactic and pragmatic completeness. Findings 1 and 4 show why that is not enough on its own: a caller saying "My account number is 4417. And the zip is..." produces a complete sentence at every digit group.
For the human timing baselines (the roughly 200 ms modal gap between speakers) and how to score them, see our guide to turn-taking evaluation.
Published benchmarks, and who published them
Only one public benchmark compares most of these methods on the same audio under the same policy: eot-bench. It is LiveKit's own benchmark, built to evaluate LiveKit's own model, released in June 2026 with open code and an open 14-language dataset of real human-to-agent turns. Treat it as a well-documented vendor benchmark, not a neutral one.
Its method is the right one, and worth copying. Each user turn is annotated with every silence of at least 100 ms. The final silence is the true end; every earlier silence is a "hold." The harness sweeps three policy knobs for every model (threshold, minimum silence before acting, and a timeout) and reports two operating points: the best false-cutoff rate at a latency budget, and the best latency at a false-cutoff budget. Latency here means dead air after the caller finished, not model inference time.

English results from the eot-bench README (lower is better):
| Model | False cut-offs at 300 ms | False cut-offs at 600 ms | Latency at 5% cut-offs | Latency at 10% cut-offs |
|---|---|---|---|---|
| LiveKit Turn Detector v1 | 9.9% | 4.5% | 543 ms | 295 ms |
| Deepgram Flux | 12.9% | 9.9% | 1,151 ms | 548 ms |
| LiveKit Turn Detector v1-mini | 27.8% | 12.1% | 1,070 ms | 698 ms |
| Smart Turn v3.2 | 35.2% | 14.8% | 1,051 ms | 739 ms |
| OpenAI GPT Realtime 2 (semantic VAD) | not reached | not reached | 1,143 ms | 824 ms |
| VAD baseline (silence only) | 55.6% | 21.7% | 1,600 ms | 1,000 ms |
The README also lists AssemblyAI, Soniox, Cartesia Ink 2, Gradium, ultraVAD, VAP and JoinIn AI Baton. The chart above shows all thirteen.
Three things the table does not tell you:
- No row is telephony. The dataset is not described as 8 kHz phone audio. Smart Turn's training data does not identify telephony either.
- Event-only APIs are scored coarsely. For Soniox, Cartesia and OpenAI semantic VAD, the adapter can only see whether the endpoint event fired, so the policy sweep has less to work with. Flux and AssemblyAI expose confidence scores and get a finer sweep.
- It measures the detector, not your pipeline. Framework delays on top of the detector are excluded. The next section shows how big those can be.
The price of halving your cut-off rate
Here is a derived number that neither the README nor any vendor page states. Subtract the 10% column from the 5% column. That is how many milliseconds of dead air each method charges to halve its premature cut-offs, the same quantity Raux and Eskenazi's exponential model predicts.
| Method | Extra latency to go from 10% to 5% cut-offs |
|---|---|
| Soniox | +135 ms |
| LiveKit Turn Detector v1 | +248 ms |
| Smart Turn v3.2 | +312 ms |
| OpenAI GPT Realtime 2 | +319 ms |
| LiveKit v1-mini | +372 ms |
| VAD baseline | +600 ms |
| Deepgram Flux | +603 ms |
Flux is competitive at a 10% budget but its curve is steep at the strict end. On this dataset, pushing Flux below 5% cut-offs costs about as much latency as a silence timer does. If your use case cannot tolerate cut-offs, such as card numbers or medication names, test Flux at `eot_threshold` 0.85 and above on your own calls before committing.
Deepgram's own numbers come from its Flux launch post and are Deepgram's claims on Deepgram's test sets: 200 to 600 ms lower response latency than pipeline approaches, about 30% fewer false interruptions, p90 end-of-turn latency of 1 s and p95 of 1.5 s, and `EagerEndOfTurn` arriving 150 to 250 ms before `EndOfTurn` at the cost of 50 to 70% more LLM calls. One trap: the launch post names the silence fallback `eot_silence_threshold_ms`. The current API parameter is `eot_timeout_ms`. Use the docs, not the blog.
Smart Turn's numbers are published per decision in its v3.2 CPU benchmark file. We walked through them, including the Spanish row, in the Pipecat Smart Turn guide.
How the pieces combine in LiveKit and Pipecat, and the stacking trap
The detector is only one stage. The framework adds its own delays, and in some modes those delays add up rather than overlap.

LiveKit: max in VAD mode, sum in STT mode
The LiveKit turn handling docs define `endpointing.min_delay` (default 0.5 s) and `max_delay` (default 3.0 s). The meaning of `min_delay` depends on the mode:
- In VAD mode, it behaves like `max(VAD silence, min_delay)`. Silero's 0.55 s silence and the 0.5 s floor overlap, so the turn closes at 0.55 s.
- In STT mode (`turn_detection="stt"`, used with Flux), it is applied after the provider's end-of-speech signal. The docs say it is "in addition to the STT provider's endpointing delay."
- With the audio turn detector, unset values default to 0.3 s and 2.5 s instead.
So the default LiveKit plus Flux setup waits for Flux to decide, then waits another 500 ms. Worked example, with illustrative numbers you should replace with your own traces: Flux fires `EndOfTurn` 400 ms after the caller stops. LiveKit adds 500 ms, so the turn commits at 900 ms. Assume 350 ms LLM time to first token and 150 ms TTS time to first byte. First audio leaves at 1,400 ms. Lower `min_delay` to 0.1 s and the same turn plays at 1,000 ms. That is 400 ms recovered with no model change. A second, subtler trap: the LiveKit Inference parameter table for Flux lists a default `eager_eot_threshold` of 0.5, while Deepgram's own default is unset. Check which one your session actually sends.
# Illustrative LiveKit Agents 1.x configs. Verify names against your installed version.
from livekit.agents import AgentSession, TurnHandlingOptions, inference
from livekit.plugins import deepgram, silero
# A. Default path: audio turn detector, VAD supplied by the session
session_a = AgentSession(
turn_handling=TurnHandlingOptions(
turn_detection=inference.TurnDetector(), # v1 on LiveKit Cloud, v1-mini elsewhere
),
# stt, llm, tts ...
)
# B. Flux decides the turn. Shrink min_delay, because in STT mode it is ADDED to Flux's signal.
session_b = AgentSession(
stt=deepgram.STTv2(model="flux-general-en", eot_threshold=0.7, eot_timeout_ms=5000),
vad=silero.VAD.load(), # still used for barge-in detection
turn_handling=TurnHandlingOptions(
turn_detection="stt",
endpointing={"min_delay": 0.1, "max_delay": 3.0},
),
)
# C. Unsupported language: silence only, tuned for slower speakers
session_c = AgentSession(
vad=silero.VAD.load(min_silence_duration=0.7, activation_threshold=0.5, sample_rate=8000),
turn_handling=TurnHandlingOptions(turn_detection="vad"),
)Two more LiveKit details change the comparison. Preemptive generation is on by default: the LLM starts as soon as a final transcript arrives, retries up to `max_retries=3` times per turn, and skips turns longer than `max_speech_duration=10.0` s. So you are already paying for some speculative tokens. And with a realtime model, leaving the model's server-side detection on while also configuring a client turn detector causes conflicts. The docs tell you to set the model's `turn_detection=None`; the Python SDK does this automatically when you configure client-side turn-taking.
Pipecat: one absolute deadline, by design
Pipecat's user turn strategies avoid the stacking problem on purpose. The STT fallback timer is an absolute deadline: end of speech plus the STT service's `ttfs_p99_latency`. VAD `stop_secs` and Smart Turn inference both fall inside that budget. The defaults are VAD plus transcription to start a turn, and `TurnAnalyzerUserTurnStopStrategy(LocalSmartTurnAnalyzerV3())` to stop it. A 5.0 s `user_turn_stop_timeout` is the backstop.
With Flux, `DeepgramFluxSTTService` requests `ExternalUserTurnStrategies` on its own, so Flux owns turn boundaries and VAD becomes optional. That external stop strategy has its own `timeout` parameter (default 0.5 s), documented as a short delay for consecutive or late transcriptions. Log the gap between Flux's `on_end_of_turn` and the bot's first audio to confirm what it costs on your pipeline. Our guide to measuring Pipecat latency shows how to instrument it.
# Illustrative Pipecat 1.x: Flux owns the turn, with eager generation enabled.
import os
from pipecat.services.deepgram.flux.stt import DeepgramFluxSTTService
stt = DeepgramFluxSTTService(
api_key=os.getenv("DEEPGRAM_API_KEY"),
enable_eager_end_of_turn=True, # switches to EagerUserTurnStrategies
settings=DeepgramFluxSTTService.Settings(
eot_threshold=0.75,
eager_eot_threshold=0.5, # must be <= eot_threshold
eot_timeout_ms=5000,
),
)
@stt.event_handler("on_end_of_turn")
async def log_eot(service, transcript):
# Write a timestamp here; compare it with the first TTS audio frame later.
...Eager and speculative generation: latency won, tokens burned
Speculative generation starts the LLM before the turn is confirmed. Flux exposes it as `EagerEndOfTurn`; LiveKit as `preemptive_generation`; Pipecat as `EagerUserTurnStrategies`, which holds the response until the turn commits and discards it if the transcript differs.
The latency gain is real but bounded by how early the eager signal fires. Deepgram cites 150 to 250 ms. The cost is extra LLM calls, and each one re-sends your whole prompt.
Worked example. Every input here is an assumption; substitute your own:
- 2,000 calls per day, 10 caller turns each: 20,000 turns per day
- 4,000 input tokens per LLM call (system prompt, tools, history)
- 60% extra calls, the middle of Deepgram's 50 to 70% range: 12,000 extra calls per day
- Cancelled drafts emit about 40 output tokens before cancellation
- Illustrative prices: $0.40 per million input tokens, $1.60 per million output tokens
Extra input: 12,000 × 4,000 = 48 million tokens per day, or $19.20. Extra output: 12,000 × 40 = 480,000 tokens, or $0.77. That is about $20 per day, or roughly $600 per 30-day month.
At an assumed 4-minute average call, you run 8,000 minutes per day. The speculative spend is about $0.0025 per minute. Flux English is listed at $0.0065 per minute pay-as-you-go (Deepgram pricing, current promotional rate). So eager mode can add a cost equal to roughly 40% of your STT bill, to buy 150 to 250 ms.
Three ways to keep that number down:
1. Use prompt caching. The re-sent prefix is identical across drafts, so cached-input pricing applies on providers that support it.
2. Draft with a smaller model, as Deepgram's eager guide suggests, and only call the full model on `EndOfTurn`.
3. Log the ratio. Count `EagerEndOfTurn` followed by `TurnResumed` versus by `EndOfTurn`. If more than half of drafts are discarded, raise `eager_eot_threshold`.
Whether 200 ms is worth $600 a month depends on your call type. In a scheduling flow where callers mostly answer yes or no, the gain is small because `EndOfTurn` fires quickly anyway. In open-ended support with tool calls, the gain is larger. Measure before and after, as described in our guide to shipping voice agent changes safely.
Decision matrix: which end-of-turn method to use

| Situation | First choice | Second choice | Why |
|---|---|---|---|
| LiveKit agent on LiveKit Cloud, English | Audio `TurnDetector` v1 (default) | Flux, `turn_detection="stt"`, lower `min_delay` | v1 is free there and led eot-bench at both budgets |
| LiveKit self-hosted | `v1-mini` on compute-optimized CPU | Flux in STT mode | v1 needs LiveKit Inference; v1-mini and Flux were close at the 5% budget in eot-bench |
| Pipecat agent | Smart Turn v3.2 (default) | Flux with `ExternalUserTurnStrategies` | Smart Turn is free and local; Flux removes a separate STT |
| Spanish and English callers in the US | Flux `flux-general-multi` with `language_hint` ["en", "es"] | LiveKit audio detector or Smart Turn | All three cover Spanish; report cut-off rate per language |
| Callers in Vietnamese, Polish, Ukrainian and others | Smart Turn v3.2 (23 languages) | Silence-only VAD with a longer timer | LiveKit and Flux do not list these languages |
| OpenAI Realtime speech-to-speech | `semantic_vad`, `eagerness: low` for intake flows | Client-side LiveKit detector with `turn_detection=None` on the model | Gemini Live does not support client-side turn-taking in LiveKit |
| Card, phone or address capture | Raise thresholds for that step only | `FilterIncompleteUserTurnStrategies` in Pipecat | Per-context thresholds were the core gain in Raux and Eskenazi |
| Target under 300 ms of dead air | Model-based detector plus speculative generation | Accept a measured cut-off rate | Even the best eot-bench result had 9.9% cut-offs at 300 ms |
| Lowest cost at scale | Silero plus Smart Turn or v1-mini on CPU | Flux if it replaces your current STT | Flux English lists at $0.0065/min vs Nova-3 at $0.0048/min (promotional rates) |
For a longer view of the endpointing landscape as a whole, see best endpointing for voice agents in 2026. For per-language STT tradeoffs, see multilingual STT for voice agents.
Changing thresholds mid-call
The research says per-context thresholds win. All three stacks let you change them without reconnecting:
# Illustrative: be patient while the caller reads digits, then go back to normal.
# LiveKit (Flux via the Deepgram plugin):
stt.update_options(eot_threshold=0.85, eot_timeout_ms=8000) # before asking for the card number
stt.update_options(eot_threshold=0.7, eot_timeout_ms=5000) # after the number is confirmed
# Pipecat (Flux):
from pipecat.frames.frames import STTUpdateSettingsFrame
await worker.queue_frame(STTUpdateSettingsFrame(
delta=DeepgramFluxSTTService.Settings(eot_threshold=0.85)))On the raw Flux socket, the same change goes out as a `Configure` control message. For yes-or-no confirmations, go the other way: lower the threshold, because Raux and Eskenazi measured the biggest latency win after confirmation prompts.
Gotchas that only show up in production
- Twilio sends 8 kHz μ-law; Flux wants linear16. Pipecat's `flux_encoding` must be `"linear16"`. Silero accepts only 8,000 or 16,000 Hz. Smart Turn takes 16 kHz. Every hop resamples, and none of the turn models has published 8 kHz results except the research paper above.
- `utterance_end_ms` is not an end-of-turn signal. Its minimum is 1,000 ms, it depends on interim results that arrive about once per second, and Deepgram documents it firing mid-thought. If your agent waits on `UtteranceEnd`, you have built a 1-second-plus timer.
- Burstable instances break CPU turn models. LiveKit warns that `v1-mini` and the text model can time out on AWS t3 or t4g instances when CPU credits run out. A timeout means the turn commits on silence alone.
- The LiveKit `v1` fallback is silent and sticky. If LiveKit Inference is unreachable or your free allowance runs out, the session emits a probability of 1.0 for the in-flight turn and switches to `v1-mini` until the session ends. Log the warning and count it.
- A higher Silero threshold does not filter TV voices. It filters quiet speech. For background talkers, use noise or voice isolation before VAD. See our post on self-hosted LiveKit noise and echo cancellation.
- Flux timeouts hide in averages. A turn closed by `eot_timeout_ms` has a `trigger` of `timeout`. Five seconds of dead air on 3% of turns barely moves a mean but ruins those calls. Track p95 and the timeout share separately.
- `eager_eot_threshold` above `eot_threshold` is an error. Deepgram rejects it, and the LiveKit plugin raises `ValueError` on `update_options`.
Our LiveKit latency debugging guide and false interruption guide cover the barge-in side, which uses the same VAD but a different decision.
How to compare end-of-turn methods on your own calls
Vendor benchmarks rank detectors on their data. You need the ranking on yours. This protocol produces one curve per method: premature cut-off rate against dead-air latency, the same axes eot-bench uses.
1. Collect 200 to 400 real caller turns from your own calls. Use stereo recordings with caller and agent on separate channels. Over-sample hard turns: digit strings, addresses, names, "let me check," and answers to open questions. Include your real phone path, so 8 kHz audio that went through your SIP provider.
2. Label every silence of 100 ms or more inside each caller turn. Run a VAD on the caller channel to find silences, then have a person mark each as "hold" (caller continued) or "end" (caller was done). The final silence before the agent should speak is the end. Our pause handling guide has labeling rules.
3. Replay each turn through each candidate method offline. Stream the caller audio in real time to Flux, Smart Turn, the LiveKit detector and a VAD timer. Record every decision with its timestamp, and the confidence where the API exposes one. Use the same audio for every method.
4. Sweep each method's knobs. For Flux, `eot_threshold` from 0.5 to 0.9. For Silero timers, silence from 0.3 s to 1.2 s. For model outputs, the probability threshold. Each setting gives one point on the curve.
5. Score each point with two numbers. Premature cut-off rate is the share of turns where the first "done" decision fell inside a hold. Latency is the time from the true end to the decision, over correctly ended turns, reported as p50 and p95. Count timeouts separately.
6. Add your framework's delay. Replay the winning configurations through the full LiveKit or Pipecat pipeline and measure first agent audio, so the stacking delays are included.
7. Pick the operating point by use case, then re-run on every change. A typical bar is under 5% premature cut-offs with p50 dead air under 700 ms, tightened to under 2% for steps that capture numbers. Re-run whenever you change STT, model version, framework version or prompts.
A minimal scorer for steps 4 and 5, assuming you have exported labeled silences and per-method decisions to JSON:
# Illustrative scorer: premature cut-off rate vs latency, one row per (method, setting).
# silences.json: [{"turn_id": "c12-t3", "start": 4.21, "end": 4.88, "label": "hold"}, ...]
# decisions.json: [{"method": "flux", "setting": 0.7, "turn_id": "c12-t3",
# "t": 4.65, "trigger": "model"}, ...] # first "done" decision per turn
import json, statistics
from collections import defaultdict
silences = json.load(open("silences.json"))
decisions = json.load(open("decisions.json"))
turn_end = {s["turn_id"]: s["start"] for s in silences if s["label"] == "end"}
holds = defaultdict(list)
for s in silences:
if s["label"] == "hold":
holds[s["turn_id"]].append((s["start"], s["end"]))
def in_hold(turn_id, t):
return any(a <= t < b for a, b in holds[turn_id])
rows = defaultdict(lambda: {"turns": 0, "premature": 0, "timeouts": 0, "lat": []})
for d in decisions:
r = rows[(d["method"], d["setting"])]
r["turns"] += 1
if d["t"] < turn_end[d["turn_id"]] - 0.05 or in_hold(d["turn_id"], d["t"]):
r["premature"] += 1 # fired before the caller was done
else:
r["lat"].append(d["t"] - turn_end[d["turn_id"]])
r["timeouts"] += d.get("trigger") == "timeout"
for (method, setting), r in sorted(rows.items()):
lat = sorted(r["lat"]) or [float("nan")]
p95 = lat[min(len(lat) - 1, int(0.95 * len(lat)))]
print(f"{method:12s} {setting:>5} cutoff={r['premature']/r['turns']:.1%} "
f"p50={statistics.median(lat)*1000:.0f}ms p95={p95*1000:.0f}ms "
f"timeouts={r['timeouts']}")The 50 ms tolerance before the labeled end mirrors the research paper's asymmetric window: a decision a hair early is fine, a decision during a real hold is not.
How many turns you need. To tell an 8% cut-off rate from a 4% one at 95% confidence and 80% power, the two-proportion formula gives n = (1.96 + 0.84)² × (0.08 × 0.92 + 0.04 × 0.96) / (0.04)² = 7.84 × 0.112 / 0.0016, or about 550 turns per method. Turns from the same caller are correlated, so add margin: 600 to 800 turns, drawn from at least 100 different calls, is a safer floor. Comparing 3% against 2% needs several thousand turns, which is why small changes need production sampling, not a test set.
Score digits and names on the transcript too, because a premature cut-off often shows up as a truncated entity rather than an obvious interruption. Our entity accuracy guide covers that scoring.
Where independent evaluation fits
Each vendor benchmark here was built by the company whose model it ranks. That does not make the numbers wrong, but it does mean none of them used your callers, your phone path or your framework settings. Independent evaluation helps in three places. Before launch, Evalgent runs a fixed set of labeled pauses, digit strings and noisy phone audio against your candidate configurations and reports the cut-off versus latency curve for each. During a bake-off, the same audio goes to every option, so Flux, Smart Turn and the LiveKit detector are compared on equal terms. After launch, scoring live calls for truncated turns and timeout-closed turns catches regressions when a model version or framework default changes underneath you.
Frequently asked questions
What is end-of-turn detection in a voice agent?
End-of-turn detection decides when a caller has finished speaking so the agent can reply. It is separate from voice activity detection, which only decides whether audio contains speech. Methods range from silence timers to STT-integrated models like Deepgram Flux and audio models like LiveKit's turn detector and Smart Turn. The goal is few premature cut-offs with little dead air.
Is Deepgram Flux better than Silero VAD for end of turn?
On LiveKit's eot-bench, Flux had far fewer cut-offs than a silence-only baseline: 12.9% versus 55.6% at a 300 ms budget. But Silero is a speech detector, not a turn model, so the fair comparison is Flux against Smart Turn or the LiveKit detector. Flux needs about 1.15 s of dead air to reach 5% cut-offs there, versus 0.54 s for LiveKit v1.
What Silero VAD threshold should I use on phone calls?
Start with the framework default: 0.5 in LiveKit, 0.7 in Pipecat. Raising it mostly drops quiet callers; it does not filter background voices, because Silero scores any human voice highly. Tune the silence duration instead, and let a turn model make the final decision. Run Silero at 8 kHz or 16 kHz only.
Does the LiveKit turn detector support multilingual callers?
Yes. The audio TurnDetector supports 14 languages, including English, Spanish, French, German, Hindi, Japanese, Korean, Arabic and Chinese, with per-language calibrated thresholds you can override through unlikely_threshold. It uses the language your STT reports, defaulting to English thresholds without STT. The older text MultilingualModel also covers 14 languages but is deprecated.
How does turn detection work with the OpenAI Realtime API?
The Realtime API detects turns server-side with server_vad, a silence timer, or semantic_vad, a classifier on the user's audio. You can disable both by setting turn_detection to null and run your own detector, such as LiveKit's. LiveKit supports this client-side mode for OpenAI Realtime but not for Gemini Live.
What is semantic VAD eagerness?
Eagerness is the one tuning knob for OpenAI's semantic_vad. It sets how long the model may wait when it thinks the user is not finished. Low, medium and high allow maximum timeouts of 8, 4 and 2 seconds. Auto equals medium. Use low for intake or dictation-style answers and high for quick yes-or-no flows.
Should I enable Deepgram Flux EagerEndOfTurn?
Enable it when your LLM step is slow, for example with tool calls or retrieval. Deepgram says eager events arrive 150 to 250 ms earlier, at the cost of 50 to 70% more LLM calls. At typical prompt sizes that can approach half your STT bill. Use prompt caching and log how often drafts are discarded.
How do I measure premature cut-offs in my voice agent?
Record calls in stereo, label every caller silence of 100 ms or more as a hold or the true end, and replay the audio through each method. A premature cut-off is a done decision that lands inside a hold. Report its rate next to p50 and p95 latency after the true end, using about 600 turns per method.
The bottom line
Model-based end-of-turn detection beats a silence timer on every published benchmark, but the right model depends on your framework, your callers' languages and the delays your stack adds on top. Pick a default from the decision matrix, remove any stacked delays, then prove the choice on your own phone audio with a cut-off versus latency curve before callers do it for you.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more