Open door for builders.
Deepgram Flux vs Nova-3 for Voice Agents: Accuracy, Turn Detection, Latency and Cost

On this page
You run a voice agent on LiveKit or Pipecat. Nova-3 is wired in, a VAD or turn detector sits on top, and callers still get cut off mid-account-number or wait a beat too long after "yes." Deepgram says Flux fixes this with turn detection built into the recognizer. The question is whether switching is worth it for your calls, your languages and your bill.
Page one for "deepgram flux vs nova 3" is Deepgram's own launch post, a feature table in the docs, and a GitHub thread where Flux got a word wrong that Nova-3 got right. None of them tells you what each latency number measures, what the switch costs at your volume, which settings quietly add half a second in your framework, or how to test the two on phone audio.
This guide does. Every price, parameter and event name below was checked against Deepgram, LiveKit and Pipecat docs and the Deepgram changelog on October 4, 2026. Where a number is mine, it is labeled as an assumption.
Flux and Nova-3 in two paragraphs
Nova-3 is Deepgram's general speech-to-text model. You stream audio to `wss://api.deepgram.com/v1/listen?model=nova-3` and get `Results` messages: interim transcripts, then finals. It does not decide when a caller has finished. You get pause signals (`speech_final` from endpointing, `UtteranceEnd` from gap detection) and your framework's VAD or turn detector makes the call. Nova-3 also runs pre-recorded, supports diarization and multichannel, and has domain variants such as `nova-3-medical` and the new `nova-3-pharma` (released September 17, 2026, per the Deepgram changelog).
Flux is what Deepgram calls conversational speech recognition. You stream to a different endpoint, `wss://api.deepgram.com/v2/listen`, with `model=flux-general-en` or `model=flux-general-multi`. Instead of interim and final results you get `TurnInfo` messages with an `event` field: `Update`, `StartOfTurn`, `EagerEndOfTurn`, `TurnResumed` and `EndOfTurn`. The quickstart is blunt: `/v1/listen` will not work with Flux, and `model=flux` is wrong. Flux is streaming only. There is no pre-recorded Flux.
This post is about the STT model choice. If your question is which end-of-turn method to use (Flux's model-integrated detection, Silero VAD with a timer, or the LiveKit turn detector), read the companion piece on end-of-turn detection with Flux, Silero and the LiveKit turn detector.
How each model decides a turn is over
The two models differ less in how they transcribe than in what they tell you about timing. That difference drives most of what follows: latency, false cutoffs, cost and code.
Nova-3: pause signals, your decision
With Nova-3, three mechanisms run on Deepgram's side, and none of them is a turn decision.
Endpointing. A VAD inside the API watches for silence. When it sees enough, it finalizes the current segment and sets `speech_final: true`. The endpointing docs say the default is 10 ms of silence. You can raise it (`endpointing=300`) or disable it. LiveKit's Inference docs list a default of 25 ms for its hosted Nova path. Either way, that is a pause detector, not a turn detector. A caller who pauses 400 ms between digit groups produces several `speech_final` segments inside one turn.
UtteranceEnd. With `interim_results=true` and `utterance_end_ms` set, Deepgram sends an `UtteranceEnd` message after a gap of that length following the last finalized word. The Utterance End docs set the minimum at 1,000 ms and warn that values under 1,000 ms gain nothing because interim results arrive about once per second. They also say it "fires based on detecting a gap even if it determines that speech is continuing after the gap," and call it "less ideal for voice agent applications."
Finalize. Your client can send `{"type": "Finalize"}` to flush pending audio into a final transcript. Pipecat's `DeepgramSTTService` does this when its VAD detects the user stopped speaking.
So in a Nova-3 pipeline the turn decision lives in your framework: Silero VAD plus a silence timer, the LiveKit turn detector, or Pipecat's Smart Turn. The STT just has to have the final words ready when that decision fires.
Flux: the recognizer emits the turn
Flux moves the decision into the model that transcribes. The state machine docs define the contract:
- `Update` messages arrive for roughly every 0.25 seconds of transcribed audio.
- `StartOfTurn` always carries a non-empty transcript. Deepgram recommends it as the barge-in trigger over an external VAD for that reason.
- `EagerEndOfTurn` fires only if you set `eager_eot_threshold`. It always has a transcript.
- `TurnResumed` can only follow an `EagerEndOfTurn`. It means "the caller kept going, cancel your draft."
- The `EndOfTurn` transcript always matches the preceding `EagerEndOfTurn` transcript. If the words changed, you get `TurnResumed` first.
- Every `EndOfTurn` carries a `trigger`: `model`, `manual` (you sent `ForceEndTurn`) or `timeout` (`eot_timeout_ms` passed).
Three parameters control it, per the end-of-turn configuration page:
| Parameter | Range | Default | What it does |
|---|---|---|---|
| `eot_threshold` | 0.5 to 1.0 | 0.7 | Confidence needed for `EndOfTurn`. 1.0 suppresses model-driven turn ends. |
| `eager_eot_threshold` | 0.3 to 0.9 | unset | Enables `EagerEndOfTurn` and `TurnResumed`. Must be less than or equal to `eot_threshold`. |
| `eot_timeout_ms` | 500 to 60,000 | 5,000 | Silence after which a turn ends regardless of confidence. |
All three can change mid-stream with a `Configure` message. Since the August 28, 2026 changelog entry, you can also override, suppress or blend Flux's turn detection per turn. If you already trust your own detector, set `eot_threshold=1.0` and send `{"type": "ForceEndTurn"}` when your detector fires. That recipe is documented in Bring Your Own Turn Detection.
Here is what an `EndOfTurn` looks like on the wire, from Deepgram's docs:
{
"type": "TurnInfo",
"event": "EndOfTurn",
"turn_index": 0,
"audio_window_start": 0.0,
"audio_window_end": 1.7,
"transcript": "Hi I need to cancel my subscription please.",
"words": [{"word": "Hi", "confidence": 0.95, "start": 0.42, "end": 0.56}],
"end_of_turn_confidence": 0.7,
"trigger": "model"
}
The practical consequence: with Nova-3, a pause inside a turn produces a final transcript and leaves the turn decision to you. With Flux, a pause inside a turn produces an `Update` (or a `TurnResumed` if eager mode guessed wrong) and the turn stays open until the model, a timeout or your `ForceEndTurn` closes it.
Feature matrix: what each model supports today
Deepgram publishes a Flux vs Nova-3 comparison. This table merges it with the Flux feature overview, the keyterm docs, the rate-limit page and the changelog, as of October 4, 2026.
| Capability | Flux | Nova-3 |
|---|---|---|
| Endpoint | `/v2/listen` | `/v1/listen` |
| Streaming / pre-recorded | Streaming only | Both |
| Turn events (StartOfTurn, EndOfTurn, Eager, Resumed) | Yes | No |
| Endpointing, UtteranceEnd, interim results | Replaced by turn events and `Update` | Yes |
| Languages | `flux-general-en`; `flux-general-multi` covers en, es, fr, de, hi, ru, pt, ja, it, nl | 45+ languages, plus `language=multi` code-switching |
| Language control | `language_hint` (multi model only; 400 error on `flux-general-en`) | `language` parameter |
| Keyterm prompting | Yes; mid-stream update via `Configure` | Yes; mid-stream update via `Configure` since Oct 2, 2026 (global endpoint only) |
| Smart formatting | No | Yes |
| Numerals | Yes; toggle mid-stream since Sep 25, 2026 | Yes |
| Redaction | Numbers only (`numbers`, `aggressive_numbers`) | Full (PCI, SSN, numbers and more) |
| Speaker diarization | No | Yes (streaming concurrency capped at 50 on pay-as-you-go, North America) |
| Multichannel | No, mono only | Yes |
| Filler words | Transcribed by default | Off by default |
| Find and replace, search | No | Yes |
| Profanity filter | Docs conflict: comparison table says no, feature overview and Pipecat say yes | Yes |
| Domain models | None | `nova-3-medical`, `nova-3-pharma` |
| Encodings (raw) | linear16, linear32, mulaw, alaw, opus, ogg-opus | linear16, mulaw, alaw, opus and more |
| Sample rates (raw) | 8000, 16000, 24000, 44100, 48000 | Any declared rate |
| Recommended chunk | 80 ms, "strongly recommended" | No hard recommendation |
| Idle timeout | 60 s, kept alive by pings | 12 s without audio or KeepAlive |
| Self-hosted, EU endpoint | Yes, yes | Yes, yes |
| Streaming concurrency, pay-as-you-go NA | 150, shared by `en` and `multi` | 150 |
Three rows change real decisions.
Filler words. Flux transcribes "um" and "uh" by default. Nova-3 does not. If your reference transcripts were made without fillers, Flux takes an insertion error for every "um," and your WER comparison is biased against it before you start. Normalize fillers out of both sides.
Redaction. Flux only redacts numbers. If your compliance design depends on Deepgram stripping card numbers or SSNs from the live transcript with category-level control, Flux does not do that today.
Diarization concurrency. The rate-limit page caps streaming diarization at 50 concurrent requests on pay-as-you-go in North America, versus 150 for plain Nova-3 streaming. Turning on `diarize` on a live agent stream to "get speaker labels for free" can cut your ceiling by two thirds.
Phone audio: 8 kHz, mu-law and what "16000 recommended" means
Most ICP3 agents take calls over Twilio, Telnyx or a SIP trunk. The caller leg is G.711 at 8 kHz, which carries speech up to about 3.4 kHz. Flux now accepts raw `mulaw` at `sample_rate=8000` on `/v2/listen`, per the quickstart's format table. Nova-3 has accepted it for years.
The Flux docs say `16000` is recommended. That does not mean you should upsample. Resampling 8 kHz audio to 16 kHz adds samples, not information; the band above 4 kHz stays empty. The recommendation describes wideband sources, like a browser microphone. For a phone agent, the honest test is the audio you actually receive.
What the frameworks do in between matters too:
- Pipecat decodes Twilio's mu-law to PCM in the frame serializer. Its `DeepgramFluxSTTService` docs list `flux_encoding` as `linear16` and say it "must be" that, so Flux sees PCM at the pipeline sample rate. See our guide to Pipecat with Twilio and Telnyx for the transport side.
- LiveKit SIP bridges the call into a room. Your agent receives PCM frames at the room's rate, but the content is still 8 kHz band-limited from the PSTN leg.
- Chunk size. Flux strongly recommends 80 ms chunks. Twilio Media Streams sends 20 ms frames. Sending 20 ms frames works, but if you buffer, buffer to 80 ms, not 250 ms. Every millisecond you hold audio before sending is a millisecond added to end-of-turn latency that no benchmark will show.
Pricing: what Flux and Nova-3 cost per minute and per month
From deepgram.com/pricing, streaming, checked October 4, 2026:
| Model | Pay-as-you-go | Growth (prepaid, $4K+/yr) | Regular price (PAYG / Growth) |
|---|---|---|---|
| Flux English | $0.0065/min (promo) | $0.0057/min (promo) | $0.0077 / $0.0065 |
| Flux Multilingual | $0.0078/min | $0.0068/min | Not marked as promo |
| Nova-3 Monolingual | $0.0048/min (promo) | $0.0042/min (promo) | $0.0077 / $0.0065 |
| Nova-3 Multilingual | $0.0058/min (promo) | $0.0050/min (promo) | $0.0092 / $0.0078 |
Streaming add-ons on the same page: keyterm prompting $0.0013/min, redaction $0.0020/min, diarization $0.0020/min, smart formatting included. Nova-3 pre-recorded is $0.0043/min.
Two things in that table are easy to miss.
At regular prices, Flux English and Nova-3 Monolingual cost the same. Both list at $0.0077 pay-as-you-go and $0.0065 on Growth. The 26% gap you see today is a "limited-time promotional rate." Do not build a model decision on it unless you lock it in a contract.
The keyterm add-on is listed without a model qualifier. The pricing page shows it as a streaming add-on. Whether it bills on Flux streams the same way it does on Nova-3 is something to confirm on your first invoice, not assume.
The language-prompting docs also say "pricing is the same as `flux-general-en`" for Flux Multilingual, while the pricing page lists $0.0078 against Flux English's regular $0.0077. Treat the pricing page as the billing source.
Monthly cost at 50k, 150k and 500k calls
Assumptions (mine, labeled): 3.0 billed STT minutes per call, because the caller channel streams for the whole call, including while the agent speaks. Pay-as-you-go rates.
| Monthly calls | STT minutes | Flux EN (promo) | Nova-3 mono (promo) | Either, regular price | Flux multi | Nova-3 multi (promo) |
|---|---|---|---|---|---|---|
| 50,000 | 150,000 | $975 | $720 | $1,155 | $1,170 | $870 |
| 150,000 | 450,000 | $2,925 | $2,160 | $3,465 | $3,510 | $2,610 |
| 500,000 | 1,500,000 | $9,750 | $7,200 | $11,550 | $11,700 | $8,700 |
Formula: `monthly = calls x minutes_per_call x rate`. At 500,000 calls the promo gap between Flux English and Nova-3 is $2,550 a month. Add keyterm prompting to Nova-3 ($0.0013/min) and Nova-3 rises to $9,150, within $600 of Flux.

The cost that is not on the STT line: eager LLM calls
Deepgram's launch post says that with `eager_eot_threshold` at 0.3 to 0.5, `EagerEndOfTurn` arrives 150 to 250 ms earlier than `EndOfTurn`, "at the cost of 50 to 70% more LLM calls." Run the numbers with an illustrative LLM cost:
- Assume 8 caller turns per 3-minute call and $0.0015 per LLM call (about 3,000 input tokens on a small model).
- Baseline LLM cost: 8 x $0.0015 = $0.012 per call.
- With 60% more calls: +4.8 x $0.0015 = $0.0072 per call, or $0.0024 per STT minute.
That $0.0024 per minute is larger than the entire promo price gap between Flux and Nova-3 ($0.0017 per minute). At 500,000 calls it is about $3,600 a month. Eager mode is often worth it, but price it as an LLM decision. LiveKit's preemptive generation, which is on by default, has a similar effect in Nova-3 pipelines, so compare like with like.
Concurrency: when the plan limit bites before the budget
Both models allow 150 concurrent streams on pay-as-you-go and 225 on Growth in North America, per the rate-limit page. Flux's pool is shared between `flux-general-en` and `flux-general-multi`.
Worked example (assumptions mine): 500,000 calls a month over 22 business days is about 22,700 calls a day. If 12% land in the peak hour, that is 2,730 calls at 3 minutes each, or 8,190 call-minutes in 60 minutes. Average concurrency in that hour is 8,190 / 60 = 136. Add 30% headroom for bursts and you need about 180 streams. That exceeds pay-as-you-go and sits under Growth. Rate limits are per project, and Deepgram explicitly forbids spreading traffic across projects to get more.
Latency: what each "finalization" number actually measures
"Flux is faster" and "Nova-3 is fast" are both true and both useless until you pin down what is being timed. Three different clocks show up in docs and dashboards.
Flux: audio-time end-of-turn latency
Deepgram's migration guide says Flux delivers "~260ms end-of-turn detection (p50 at defaults)." The launch post defines its latency metric as "the median audio time elapsed after the user has finished speaking in the case of a successful detection," and reports a p90 of 1 second and a p95 of 1.5 seconds.
Read that definition closely:
- Audio time, not wall-clock. It excludes your network round trip, your chunk buffering and any framework delay after the event lands.
- Successful detections only. A false cutoff mid-sentence is not counted as a fast detection; it is counted against precision.
- At defaults. Raise `eot_threshold` to 0.85 and the median moves out. Lower it and false cutoffs rise.
The p90 and p95 figures matter more than the median for a phone agent. One turn in ten taking a second or more to commit is what callers notice.
Nova-3: transcript finalization plus your turn detector
For Nova-3 there are two separate delays. First, how fast a final transcript arrives after the caller stops (endpointing setting plus processing plus network). Our post on Deepgram STT latency, diarization and stability covers how to measure that one. Second, how long your turn detector waits before declaring the turn over. In a Nova-3 pipeline the second number usually dominates.
Pipecat makes this explicit: `DeepgramSTTService` takes a `ttfs_p99_latency` parameter, described as "P99 latency from speech end to final transcript in seconds," which the framework uses to wait for the final transcript after its VAD fires. If you set it lower than reality, the turn closes before the last words arrive.
The LiveKit setting that adds 500 ms to Flux
This is the most expensive default in this comparison. LiveKit's docs tell you to use Flux's turn detection by setting `turn_detection="stt"`. The turn handling reference says that `endpointing.min_delay` defaults to 0.5 seconds and that "in STT mode, this is applied after the STT end-of-speech signal, and therefore in addition to the STT provider's endpointing delay."
So a stock LiveKit agent with Flux in STT mode waits for Flux's `EndOfTurn` and then waits another 500 ms. The audio turn detector path, by contrast, defaults `min_delay` to 0.3 seconds. Lower `min_delay` when you hand turn detection to Flux, then measure. If you skip this, your bake-off will show Flux as slower than it is.
A per-hop budget for one turn
Illustrative numbers (assumptions, except where attributed), measured from the moment the caller stops speaking to the moment the LLM request starts:
| Hop | Nova-3 + LiveKit audio turn detector | Flux in LiveKit STT mode, stock | Flux in STT mode, min_delay lowered |
|---|---|---|---|
| Chunk buffering | 10 ms | 40 ms (80 ms chunks, half on average) | 40 ms |
| Network to Deepgram and back | 40 ms | 40 ms | 40 ms |
| STT end signal | 25 ms endpointing + ~150 ms final | ~260 ms p50 (Deepgram) | ~260 ms p50 |
| Turn detector inference | ~50 ms | none | none |
| Framework min_delay | 300 ms default | 500 ms default | 100 ms (your choice) |
| Total to LLM start, p50 | ~575 ms | ~840 ms | ~440 ms |
The point is not the totals; your numbers will differ. The point is that the framework timer is the largest single hop in two of three columns, and none of the vendor latency charts include it. Measure it per hop, as in our guide to measuring Pipecat latency.
Accuracy: what the published numbers say and do not say
Vendor claims, attributed
- Deepgram's launch post says Flux "performed nearly the same across word error rate (WER) and word recognition rate (WRR) as Nova-3" on its benchmark sets and achieved the lowest WER among evaluated models on conversational audio. It reports that WER stays stable across `eot_threshold` settings, with slight degradation at the most aggressive thresholds.
- The same post claims Flux cuts agent response latency by 200 to 600 ms versus pipeline approaches and reduces false interruptions by about 30%. The test sets are Deepgram's and are not published.
- For Flux Multilingual, Deepgram's benchmark announcement reports production-audio WER of 9.94% English, 11.29% Spanish, 10.66% German, 13.62% French, 13.16% Portuguese and 17.18% Hindi, each the lowest in its chart, using "each vendor's default streaming configuration."
These are vendor benchmarks on vendor-chosen data. They are useful for ruling out a model that is far off. They do not tell you which model wins on your callers.
Independent numbers
Artificial Analysis lists Nova-3 at 5.2% on its AA-WER index. Two caveats limit how far that carries. The index is non-streaming, so it measures Nova-3 in batch mode. And Flux, which has no batch mode, is not on the leaderboard. As of this writing there is no independent, published Flux vs Nova-3 comparison on streaming phone audio.
The GitHub thread, and what it teaches
In Deepgram discussion #1497, a team reported Flux transcribing "admin console" as "ad pain console" where Nova-3 got it right. Deepgram's reply: the comparison was Flux streaming against Nova-3 batch, and "Nova-3 in batch mode has full utterance context, while Flux streaming has to make low-latency, incremental decisions," so "we don't expect full parity."
That reply is the most useful sentence in this whole comparison. If your post-call transcripts come from Nova-3 pre-recorded and your live agent runs Flux, you will see disagreements, and some will be Flux errors that a batch model would not make. Compare streaming to streaming. Use batch only as a reference transcript, and even then, verify by hand.
Why your own 8 kHz calls decide it
Vendor and leaderboard tests skew toward wideband audio and general vocabulary. Your agent hears band-limited phone audio, speakerphones, car noise, and the names, numbers and SKUs your business runs on. Our guides to STT evaluation for voice agents and entity accuracy explain why a model can win on WER and lose on the eight digits that matter.
Two research findings sharpen this for Flux vs Nova-3.
Pauses inside a turn are longer than gaps between turns. A 2026 paper on turn-aware streaming ASR (Li and Shi, arXiv 2609.04225) builds its case on prior conversation research: humans exchange turns with median gaps near 200 ms, deployed silence timeouts sit at 500 to 1,000 ms, and within-turn pauses routinely run longer than between-turn gaps. Its Figure 1 example is a caller reading a phone number: a 0.7 s timeout fires inside the number, a 1.0 s timeout leaves "I'm still here" hanging. This is exactly where a turn-aware model like Flux should beat a silence timer on top of Nova-3, and exactly where you should test it: dictated digits.
Biasing can inject words that were not said. The same paper found that when a context prefix always matched the audio during training, the model learned to copy it: 40% of wrong-profile probes pulled the wrong entity into the transcript, cut to 0.8% with counterfactual training. Keyterm prompting is a different mechanism, but the failure to test for is the same. Measure keyterm false insertions: how often a boosted term appears when the caller said something else. Deepgram's keyterm docs also note that malformed keyterms (commas, semicolons, `term:0.15` weights) "silently boost nothing" rather than erroring, so check the URL you actually send.
Normalize before you compare. WER is sensitive to formatting: "$1,500" vs "fifteen hundred dollars," "Dr." vs "doctor," fillers, casing. The Whisper authors released an English text normalizer for this reason (Radford et al., arXiv 2212.04356). Nova-3 with `smart_format` and Flux without it will format numbers differently; score both after the same normalization, and score entities separately. Our word error rate guide covers the normalization rules.
Integration: Flux and Nova-3 in LiveKit and Pipecat
LiveKit Agents
Flux uses the `STTv2` class, which connects to `/v2/listen`. Nova-3 uses `STT`, which connects to `/v1/listen`. From the LiveKit Deepgram docs (simplified):
from livekit.agents import AgentSession, TurnHandlingOptions
from livekit.plugins import deepgram
# Flux: let the STT own end-of-turn, and stop LiveKit from adding 500 ms after it
flux_session = AgentSession(
stt=deepgram.STTv2(
model="flux-general-en",
eot_threshold=0.7,
eager_eot_threshold=0.5, # must be <= eot_threshold
eot_timeout_ms=5000,
keyterm=["Acme", "PolicyMax"],
),
turn_handling=TurnHandlingOptions(
turn_detection="stt",
endpointing={"mode": "fixed", "min_delay": 0.1, "max_delay": 3.0},
),
# ... llm, tts, vad
)
# Nova-3: Deepgram reports pauses, LiveKit's turn detector decides
nova_session = AgentSession(
stt=deepgram.STT(model="nova-3", language="en-US", keyterm=["Acme", "PolicyMax"]),
# turn_handling omitted: defaults to LiveKit's audio TurnDetector
# ... llm, tts, vad
)Mid-call changes on Flux go through `stt.update_options(...)`. LiveKit applies `eager_eot_threshold`, `eot_threshold`, `eot_timeout_ms`, `keyterm` and `language_hint` over the open connection without interrupting the turn. Changing `model`, `sample_rate`, `mip_opt_out` or `tags` reconnects the stream.
Two LiveKit-specific gotchas:
- Threshold ordering. In livekit/agents #5199, `eager_eot_threshold=0.8` with the default `eot_threshold` (0.7) failed the WebSocket handshake with a bare HTTP 400. The current docs say the plugin now raises a `ValueError` on update, but check your version.
- Range mismatch. LiveKit documents `eot_threshold` as 0.5 to 0.9. Deepgram documents 0.5 to 1.0, where 1.0 is the "bring your own turn detection" switch. If you want that mode, confirm your plugin version accepts 1.0.
Pipecat
Flux is `DeepgramFluxSTTService`; Nova-3 is `DeepgramSTTService`. The `InputParams` and `live_options` patterns were deprecated in v0.0.105 in favor of `Settings`. From the Pipecat Deepgram docs (simplified):
import os
from pipecat.services.deepgram.flux.stt import DeepgramFluxSTTService
from pipecat.services.deepgram.stt import DeepgramSTTService
from pipecat.frames.frames import STTUpdateSettingsFrame
flux_stt = DeepgramFluxSTTService(
api_key=os.getenv("DEEPGRAM_API_KEY"),
enable_eager_end_of_turn=True, # speculative reply, held until the turn is confirmed
settings=DeepgramFluxSTTService.Settings(
model="flux-general-en",
eot_threshold=0.7,
eager_eot_threshold=0.5,
eot_timeout_ms=5000,
keyterm=["Acme", "PolicyMax"],
),
)
nova_stt = DeepgramSTTService(
api_key=os.getenv("DEEPGRAM_API_KEY"),
settings=DeepgramSTTService.Settings(
model="nova-3-general",
language="en",
smart_format=True,
keyterm=["Acme", "PolicyMax"],
),
)
# During digit collection, make Flux more patient (sent as a Configure message, no reconnect)
async def enter_digit_collection(worker):
await worker.queue_frame(STTUpdateSettingsFrame(
delta=DeepgramFluxSTTService.Settings(eot_threshold=0.85, eot_timeout_ms=7000)
))What the Pipecat docs say about Flux that changes your pipeline:
- Flux requests `ExternalUserTurnStrategies` automatically (or `EagerUserTurnStrategies` with eager mode), so you do not configure turn strategies by hand.
- A VAD is optional with Flux. Keep one if you want useful STT metrics.
- `keyterm`, thresholds, `profanity_filter` and `language_hints` update over `Configure`. `model`, `numerals` and `redact` force a reconnect, deferred until the user stops speaking.
- Flux emits no interim transcriptions. Subscribe to `on_update` if you need incremental text.
One production report deserves attention before you raise thresholds. In pipecat #5735, a team on Telnyx at 8 kHz with `eot_threshold=0.85` and `eot_timeout_ms=7000` had background noise open a turn that produced no `EndOfTurn` for 23 seconds while the caller spoke and then hung up. Their direct replay against Flux found `EndOfTurn` withheld while `Update` messages continued, with peak `end_of_turn_confidence` at 0.8208, just under their 0.85 threshold. At 0.7 / 5000 the same audio closed normally. The thread also notes Pipecat's local turn stop timeout is VAD-gated, and that `ForceEndTurn` was not exposed in v1.10.0. Lesson: a high `eot_threshold` plus noise can strand a turn. Cap turn duration in your own code and test noise-opened turns explicitly.
Gotchas collected from docs, changelogs and issues
1. Wrong endpoint, silent confusion. Flux needs `/v2/listen`. Sending `language=en` to Flux is wrong; the model name selects language. `language_hint` on `flux-general-en` returns 400.
2. Promo pricing ends. The Flux vs Nova-3 price gap exists only at promotional rates.
3. Keyterms that do nothing. Comma-separated or weighted keyterms are accepted and ignored. Repeat the parameter instead.
4. Mid-stream keyterm updates on Nova-3 are global-endpoint only. The October 2 changelog says EU, Australia and India endpoints do not support them yet.
5. Nova-3 drops idle sockets after 12 seconds without audio or `KeepAlive`. Hold music and warm transfers trip this. Flux's timeout is 60 seconds.
6. UtteranceEnd fires on gaps even when speech continues. Do not use it alone as a turn signal.
7. `last_word_end: -1` in an `UtteranceEnd` means the result was already finalized. Ignore it.
8. Eager mode doubles your LLM traffic in the worst case. Budget the tokens, and never fire irreversible tool calls on `EagerEndOfTurn`.
9. Diarization lowers streaming concurrency to 50 on pay-as-you-go.
10. Flux TTS is a different product. Deepgram also sells Flux TTS on `/v2/speak` at $0.045 per 1,000 characters. Searching "deepgram flux tts" will mix the two.
Decision matrix: Flux or Nova-3 by use case
| Your situation | Lean | Why | What to verify in the bake-off |
|---|---|---|---|
| English phone support, short turns, frequent barge-in | Flux EN | `StartOfTurn` with guaranteed transcript for barge-in; turn decision inside the model | False cutoff rate, p90 end-of-turn latency, barge-in stop time |
| Entity-heavy: account numbers, dictated digits, emails | Test both | Flux's turn awareness helps digit pauses; Nova-3 has smart formatting and find-and-replace | Entity error rate per entity type, premature cutoffs inside digit strings |
| Pharmacy or clinical vocabulary | Nova-3 domain model | `nova-3-medical` and `nova-3-pharma` exist; Flux has no domain variants | Drug and term error rate with and without keyterms |
| Spanish/English or other languages in Flux's 10 | Flux multi | `language_hint`, detect-then-lock via `Configure`, native code-switching | Per-language WER, language detection on first turn |
| Languages outside Flux's 10 (Vietnamese, Korean, Arabic, Tagalog) | Nova-3 | Flux does not support them | Not applicable |
| Need diarization, multichannel, PCI or SSN redaction | Nova-3 | Flux lacks these | Concurrency under diarization |
| You already tuned a turn detector you trust | Either | Flux with `eot_threshold=1.0` plus `ForceEndTurn`, or keep Nova-3 | Transcript accuracy only; turn timing stays yours |
| High volume, turn-taking already fine, cost first | Nova-3 during promo | $0.0017/min cheaper today; equal at list | Whether the gap survives your contract |
| Post-call transcripts and analytics | Nova-3 pre-recorded | Flux is streaming only | Agreement between live and batch transcripts |

For a broader view of how accuracy trades against price across providers, see STT cost vs accuracy for voice agents.
How to run a Flux vs Nova-3 bake-off on your own calls
This protocol takes two to four days for one engineer. It compares the two models on the same audio, streamed the same way, scored the same way.
1. Build the dataset from real calls. Pull 200 to 300 caller-channel recordings from production at 8 kHz mu-law, with consent. Stratify: 40% clean, 30% noisy (car, street, TV), 20% entity-heavy turns (account numbers, dates, names, emails), 10% accented or non-native speech. If you serve Spanish, add a separate Spanish stratum. Our guide to benchmarking voice agents on your own data covers sampling and consent.
2. Label two things per turn. A verbatim reference transcript, and turn boundaries: `speech_start`, `speech_end` and any within-turn pauses over 300 ms. Mark every entity span with its type and canonical value.
3. Stream, do not upload. Send each file to both endpoints at real-time pace in 80 ms chunks, with the same encoding the production path uses. Batch Nova-3 is not a fair comparison, as discussion #1497 shows.
4. Fix configs before you look at results. Flux: defaults (`eot_threshold=0.7`), plus one tuned variant. Nova-3: your production settings (`endpointing`, `utterance_end_ms`, `keyterm`, `smart_format`). Use the same keyterm list on both. Write the configs down.
5. Record every message with arrival time. Flux `TurnInfo` events; Nova-3 `Results` with `is_final` and `speech_final`, plus `UtteranceEnd`.
6. Score transcript accuracy. Normalize both hypothesis and reference (lowercase, strip punctuation, drop fillers, spell out or digitize numbers consistently), then compute WER. Separately compute entity error rate: the share of labeled entities not reproduced exactly after normalization.
7. Score turn timing. For Flux: end-of-turn latency (`EndOfTurn` arrival minus labeled `speech_end`) and false cutoffs (`EndOfTurn` before `speech_end`). For Nova-3: `speech_final` and `UtteranceEnd` timing, then replay through your real turn detector if you want an apples-to-apples turn comparison.
8. Check for keyterm false insertions. Count boosted terms in hypotheses where the reference does not contain them.
9. Decide with thresholds set in advance. For example: switch to Flux only if entity error rate is no worse by more than 1 point, false cutoff rate falls, and p90 end-of-turn latency is under 1,000 ms wall-clock.
Streaming harness (illustrative)
This harness streams one mu-law file to both endpoints and logs arrival times. It uses the raw WebSocket APIs documented by Deepgram, so it does not depend on SDK versions.
# bakeoff_stream.py -- illustrative, simplified
import asyncio, json, os, time
import websockets # websockets >= 14 for additional_headers
KEY = os.environ["DEEPGRAM_API_KEY"]
HDR = {"Authorization": f"Token {KEY}"}
CHUNK_MS = 80
BYTES_PER_CHUNK = 8000 * CHUNK_MS // 1000 # mu-law: 1 byte per sample at 8 kHz -> 640
FLUX_URL = ("wss://api.deepgram.com/v2/listen?model=flux-general-en"
"&encoding=mulaw&sample_rate=8000&eot_threshold=0.7"
"&keyterm=Acme&keyterm=PolicyMax")
NOVA_URL = ("wss://api.deepgram.com/v1/listen?model=nova-3&language=en-US"
"&encoding=mulaw&sample_rate=8000&interim_results=true"
"&endpointing=300&utterance_end_ms=1000&smart_format=true"
"&keyterm=Acme&keyterm=PolicyMax")
async def run(url, audio: bytes, label: str):
events = []
async with websockets.connect(url, additional_headers=HDR) as ws:
t0 = time.monotonic()
async def sender():
for i in range(0, len(audio), BYTES_PER_CHUNK):
await ws.send(audio[i:i + BYTES_PER_CHUNK])
# pace to real time against the stream clock
target = t0 + (i + BYTES_PER_CHUNK) / 8000
await asyncio.sleep(max(0, target - time.monotonic()))
await asyncio.sleep(2.0) # let trailing events arrive
await ws.send(json.dumps({"type": "CloseStream"}))
async def receiver():
async for raw in ws:
msg = json.loads(raw)
events.append({"t": time.monotonic() - t0, "msg": msg})
await asyncio.gather(sender(), receiver(), return_exceptions=True)
return {"model": label, "events": events}
async def main(path):
audio = open(path, "rb").read() # raw mu-law, no WAV header
flux, nova = await asyncio.gather(run(FLUX_URL, audio, "flux"),
run(NOVA_URL, audio, "nova3"))
json.dump([flux, nova], open(path + ".events.json", "w"))
if __name__ == "__main__":
import sys; asyncio.run(main(sys.argv[1]))Scoring (illustrative)
# bakeoff_score.py -- illustrative, simplified
import re, jiwer
FILLERS = {"um", "uh", "erm", "hmm", "mm"}
def norm(text: str) -> str:
t = re.sub(r"[^\w\s']", " ", text.lower())
return " ".join(w for w in t.split() if w not in FILLERS)
def wer(ref: str, hyp: str) -> float:
return jiwer.wer(norm(ref), norm(hyp))
def entity_error_rate(entities, hyp: str) -> float:
# entities: [{"type": "account_number", "value": "4 4 7 1 9 0 2 2"}, ...]
h = norm(hyp).replace(" ", "")
missed = sum(1 for e in entities if norm(e["value"]).replace(" ", "") not in h)
return missed / max(1, len(entities))
def flux_turn_timing(events, speech_end_s):
"""Return (latency_s, false_cutoff) for the first EndOfTurn after speech_start."""
for e in events:
m = e["msg"]
if m.get("type") == "TurnInfo" and m.get("event") == "EndOfTurn":
return e["t"] - speech_end_s, e["t"] < speech_end_s
return None, False
def flux_transcript(events):
return " ".join(e["msg"]["transcript"] for e in events
if e["msg"].get("event") == "EndOfTurn")
def nova_transcript(events):
return " ".join(e["msg"]["channel"]["alternatives"][0]["transcript"]
for e in events
if e["msg"].get("type") == "Results" and e["msg"].get("is_final"))Report medians and p90s per stratum, not one blended number. A model that wins on clean audio and loses on noisy entity turns is the common result, and the blended average hides it.
How many calls you need
For a paired comparison of entity accuracy, assume the two models disagree on 8% of entities and you want to detect a 3-point difference at 80% power and 5% significance. A McNemar approximation gives roughly `n = (1.96 x sqrt(0.08) + 0.84 x sqrt(0.08 - 0.03^2))^2 / 0.03^2`, about 700 entities. At 3 entities per entity-heavy call, that is about 235 calls in that stratum alone. If you only have 60 entity-heavy calls, you can detect large differences but not a 3-point one, and you should say so in the decision memo.
Testing the choice after you make it
A bake-off answers "which model today." It does not catch the regressions that come later: a Deepgram model update, a new keyterm list, an `eot_threshold` change someone pushed to fix one complaint, or a framework upgrade that changes a default. Treat the bake-off set as a regression suite.
- Before every config change, replay the set through both the old and new config and compare entity error rate, false cutoffs and p90 end-of-turn latency against fixed thresholds.
- In production, log every Flux `EndOfTurn` with its `trigger` and `end_of_turn_confidence`. A rising share of `timeout` triggers means turns are ending on the backstop, not the model. Count turns longer than 20 seconds; that is the stranded-turn failure from pipecat #5735.
- Weekly, sample calls and score transcripts against human references, with entities scored separately. Our Deepgram STT testing guide has a full test matrix.
This is where an independent evaluator helps a team without an eval function. Evalgent runs vendor bake-offs on your recorded calls, with the same audio and scoring for every model, and keeps the set as a regression gate so a threshold change gets measured before it reaches callers. The value of independent voice AI evaluation here is simple: the team that picked the model is not the team grading it. For multilingual lines, the same approach applies per language; see multilingual STT for voice agents.
Frequently asked questions
Is Deepgram Flux more accurate than Nova-3?
Deepgram reports near-identical WER for Flux and Nova-3 streaming on its benchmarks. No independent streaming comparison is published. Community reports show Flux streaming losing to Nova-3 batch on individual words, which Deepgram attributes to batch having full context. On your calls, either can win; entity-heavy and noisy turns are where they diverge. Test both on recorded 8 kHz audio.
What is the difference between Flux end-of-turn and Nova-3 endpointing?
Nova-3 endpointing is a silence detector: after a set pause it finalizes a segment and sets speech_final. It does not know whether the caller is done. Flux scores end-of-turn confidence from words and audio together and emits EndOfTurn when confidence passes eot_threshold, or when eot_timeout_ms of silence passes, or when you send ForceEndTurn.
How much does Deepgram Flux cost?
On October 4, 2026, Flux English streams at $0.0065 per minute pay-as-you-go and $0.0057 on Growth, both promotional. Regular prices are $0.0077 and $0.0065, the same as Nova-3 Monolingual. Flux Multilingual is $0.0078 and $0.0068. Add-ons such as keyterm prompting are priced separately; confirm on your invoice how they apply to Flux.
Which languages does Flux Multilingual support?
flux-general-multi supports English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch. Pass language_hint to bias toward expected languages, or omit it to auto-detect. TurnInfo events report detected languages, so you can lock to one language mid-call with a Configure message. Nova-3 covers 45+ languages if yours is not on that list.
Does Flux work with Twilio 8 kHz mu-law audio?
Yes. Flux's documented formats include raw mulaw at sample_rate 8000 on /v2/listen. Upsampling to 16 kHz adds no information, so stream what you receive. Pipecat decodes Twilio audio to PCM before the Flux service, and LiveKit hands the agent PCM frames. Send 80 ms chunks if you buffer at all.
How do I use Deepgram Flux in LiveKit?
Use deepgram.STTv2 with model flux-general-en or flux-general-multi, and set turn_detection to "stt" in TurnHandlingOptions so Flux owns end-of-turn. Lower endpointing min_delay, which defaults to 0.5 seconds and is added after the STT's end signal in STT mode. Keep eager_eot_threshold at or below eot_threshold.
How do I use Deepgram Flux in Pipecat?
Use DeepgramFluxSTTService with settings=DeepgramFluxSTTService.Settings(...). It requests external user turn strategies automatically, so Flux drives turns and a VAD becomes optional. Set enable_eager_end_of_turn=True for speculative replies. Update thresholds and keyterms mid-call with STTUpdateSettingsFrame; model, numerals and redact changes reconnect the stream.
Should I switch from Nova-3 to Flux?
Switch if your complaints are about turn-taking (cutoffs, slow replies, barge-in) and your languages and features are covered. Stay on Nova-3 if you need diarization, multichannel, full redaction, a domain model, or a language outside Flux's ten. In both cases, run a streaming bake-off on your own calls with thresholds decided in advance.
The bottom line
Flux is the better fit when turn-taking is your problem and its ten languages and narrower feature set cover your calls, while Nova-3 stays the safer choice for breadth, compliance features and domain vocabulary at the same list price. Make the call with a streaming bake-off on your own 8 kHz recordings, and keep that set as the regression gate for every threshold and model change afterward.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more