Open door for builders.
Fixing False Interruptions in LiveKit Agents: Adaptive Interruption, Thresholds and How to Test

On this page
Your agent is halfway through reading a refund policy. The caller says "mm-hm." The agent stops, waits two seconds in silence, then picks up mid-sentence. Or worse, it throws the sentence away and answers a question nobody asked.
That is a false interruption. It is the most common turn-taking complaint on LiveKit Agents, and page one of Google for it is docs pages and GitHub issues. The docs list the knobs. They don't tell you which knob actually fires in which mode, what the defaults do on a phone line, or how to prove a change helped.
This guide works from the current `livekit/agents` source as well as the docs. It traces one overlap through the pipeline, from VAD to playout pause to context truncation. It covers the defaults that matter: a 3-second window at the start of inbound calls where caller speech is dropped, a `min_words` setting that discards short answers, and a silent fallback from adaptive to VAD mode. It also covers what the backchannel research says about when callers will talk over your agent. Then it gives you a cause-to-fix table, a tuning procedure, a test set with pass thresholds, and Python that logs interruption events and computes the rates.
What counts as an interruption in LiveKit Agents
LiveKit uses "interruption" for one specific event: user audio arrives while the agent is speaking, and the framework decides the agent should yield. It's a different decision from end-of-turn detection, and teams mix the two up all the time.
- The turn detector (`inference.TurnDetector()`, the default `turn_detection`) decides when the user has finished, so the agent can reply. It works on user speech while the agent is silent.
- Interruption handling decides whether the agent should stop while it's talking. It works on overlapping speech.
If the complaint is "the agent cuts me off when I pause," that's endpointing. Our turn-taking evaluation guide covers it. If the complaint is "the agent stops talking when I say 'yeah'," that's interruption handling, and that's this guide.
The InterruptionOptions reference lists the settings. In Python they all live under `turn_handling["interruption"]`, in seconds:
| Option | Default | What it gates |
|---|---|---|
| `enabled` | `True` | Master switch. The old bool shorthand is gone; use `{"enabled": False}` |
| `mode` | `"adaptive"` if available, else `"vad"` | Barge-in model or raw VAD |
| `min_duration` | 0.5 s | Speech length before VAD can pause the agent |
| `min_words` | 0 | Transcribed words required (needs STT) |
| `false_interruption_timeout` | 2.0 s | Silence after a pause before it counts as false; `None` disables |
| `resume_false_interruption` | `True` | Resume the paused speech after a false interruption |
| `backchannel_boundary` | `(1.0, 1.0)` | Python only. Start and end cooldown around each agent turn |
| `discard_audio_if_uninterruptible` | `True` | Drop user audio during uninterruptible speech |
Two things aren't in that table but shape every result. First, `aec_warmup_duration` lives on `AgentSession`, not under `turn_handling`. Second, with a realtime model that uses server-side turn detection, almost none of these apply. The turns overview says only `enabled` and `discard_audio_if_uninterruptible` still matter, and setting `enabled=False` raises a `ValueError` at session start.
The interruption pipeline, step by step
Here's what happens to one overlap in the current source (`voice/agent_activity.py`, `voice/audio_recognition.py` and `inference/interruption.py`). Knowing the order explains almost every symptom later in this guide.

1. AEC warmup can make the caller inaudible
Every user audio frame passes through `push_audio`. If the agent is speaking and the AEC warmup timer is running, the frame goes to STT as silence. The real frame still reaches VAD, but `_interrupt_by_audio_activity` returns early while warmup is active.
The session source sets the default to 3.0 seconds and describes it as "the duration in seconds that the agent will ignore user's audio interruptions after the agent starts speaking." The timer is one-shot and starts on the session's first speaking turn, which is usually the greeting. When a participant joins, the session sets the duration to `None` for outbound SIP calls and back to 3.0 s for everything else, unless you set it explicitly.
That means on inbound calls and WebRTC sessions, the first 3 seconds of the greeting are deaf. A caller who says "Hi, I need to move my Thursday appointment" over the greeting doesn't interrupt. Their words also never reach STT, because STT got silence. The events reference confirms this from the other side: `user_transcription_timeout` can fire "during acoustic echo cancellation (AEC) warmup."
Outbound calls have the opposite default. The greeting can be interrupted from the first frame, so a callee's "Hello? Hello?" overlapping your opener will stop it.
2. Two gates, depending on mode
In VAD mode, each VAD inference result checks `speech_duration >= min_duration`. If it passes and `min_words > 0`, the session also counts words in the current interim transcript. Only then does it pause the agent. A 0.3 s "mm-hm" never pauses the agent. A 0.7 s "yeah, okay, sure" does.
In adaptive mode, VAD-driven pausing is switched off for the overlap. Instead, the audio goes to LiveKit's barge-in model, and the agent pauses only if the model returns `is_interruption=True`. The exception is the start cooldown. If the overlap begins within `backchannel_boundary[0]` (1.0 s) of the agent starting to speak, the source hands the rest of that agent-speech interval back to VAD and the `min_duration` gate. So in adaptive mode, `min_duration` mostly matters in the first second of each agent turn and after a fallback.
The model client shows how the model sees audio. It streams 16 kHz mono PCM over a websocket. During overlap it sends a window every 100 ms, with a 1.0 s audio prefix and at most 3 s of audio per request. The decision threshold is set on the server unless you override it. LiveKit's launch post reports inference in 30 ms or less and a median of 216 ms of audio needed to trigger an interruption.
3. The agent pauses before it decides
When a gate passes and `resume_false_interruption` is on, the agent doesn't cancel its speech. It calls `audio_output.pause()`, sets the agent state to `listening`, and keeps the speech handle alive. The caller hears the agent stop mid-word.
4. The false-interruption timer
The timer arms when VAD reports end of speech. It doesn't start at the pause. If a final transcript arrives first, the paused speech is interrupted for real and a new turn is committed. If `false_interruption_timeout` passes with no transcript, the session emits `agent_false_interruption` with `resumed=True` and continues from where it stopped.
5. Truncation
On a real interruption, the turns overview says the agent "automatically truncates its conversation history to include only the portion of the speech that the user heard." The assistant message arrives in `conversation_item_added` with `interrupted=True`. A resumed false interruption doesn't truncate, because the speech plays out in full.
Where backchannels go in adaptive mode
When the model judges an overlap to be a backchannel, the STT pipeline drops only the backchannel part of the transcript and keeps the rest of the turn. With a realtime model and client-side turn-taking, the adaptive interruption docs describe an all-or-nothing gate. A confirmed backchannel drops the whole buffered user turn and clears its audio.
Why "mm-hm," coughs, TV and echo trigger false interruptions
Every false interruption is VAD hearing speech-like energy while the agent talks. The sources differ in when they happen and which fix works.
Backchannels are predictable, and your agent invites them
Backchannels are not rare noise. Ward and Tsukahara (Journal of Pragmatics, 2000) cite an analysis of 1,155 Switchboard conversations where 19% of utterances were backchannels: 37,096 of about 205,000. Their main finding is about timing. In English, listeners tend to produce feedback after a region of low pitch, below the speaker's 26th-percentile pitch level, lasting at least 110 ms, after at least 700 ms of speech, with a reaction lag of about 700 ms. As a predictor, the rule was right 18% of the time against 13% for random placement, and it covered 48% of backchannels. That's a weak cue, but it's a real one.
The production meaning is direct. Your TTS voice produces phrase-final pitch falls all the time. Every long agent turn, such as a policy readout, a list of time slots or an order recap, creates many backchannel openings, each followed by a caller "mm-hm" about 0.7 s later. Expected false interruptions per call scale with agent talk time, not with caller behavior. Two practical rules follow:
- Shorter agent turns cut exposure. Splitting a 25-second recap into three confirmations removes most of the mid-turn openings.
- Place test backchannels the way humans do. Put them about 700 ms after an agent phrase boundary, not at random offsets. Random offsets land mid-word and understate the problem.
Voice Activity Projection (Ekstedt and Skantze, Interspeech 2022) is the research line that models this directly. It's a self-supervised model that predicts upcoming voice activity for both speakers, and it was evaluated zero-shot on predicting turn shifts and backchannels. The useful idea for builders is that backchannels are part of a predictable rhythm, not an error to filter. LiveKit's source carries an internal `agent_backchannel_opportunity` event from the turn detector, though it's marked as not yet public.
Coughs, sighs and throat clearing
These are short, high-energy and often longer than 0.5 s. In VAD mode they pass `min_duration` easily. They rarely produce a transcript, so they nearly always end in a resume. The cost is the stall, which the worked example below puts at about 2.6 seconds.
TV, radio and side conversations
These are the hardest case. Background speech is real speech, so STT transcribes it, the final transcript arrives, and the false interruption becomes a hard one. The agent drops its speech and answers the TV.
Full-Duplex-Bench v1.5 (Lin et al., ICASSP 2026) built test audio for exactly this, and its recipes are worth copying. "Talking to others" audio is cut by 8 dB, has frequencies above 4 kHz cut by 5 dB, and gets reflections at 45 ms (-6 dB) and 120 ms (-12 dB). Background speech uses a different voice, cut by 15 dB, low-pass filtered at 3 kHz, with a 100 ms echo at -10 dB.
Speakerphone and in-car echo
On a phone call, the agent's own voice comes out of the caller's speaker and can leak back into their microphone. If the handset's echo canceller doesn't remove it, VAD hears speech whenever the agent talks. You get self-interruptions, often in the first seconds of a call, before the far-end canceller has adapted. That's the problem `aec_warmup_duration` was added for. The source docstring says it is "useful to prevent the agent from being interrupted by echo before AEC is ready."
Echo cancellation and noise cancellation: what each one actually removes
The noise and echo cancellation docs draw a line that matters here. Browser `echoCancellation` and `noiseSuppression` run in the client, so they help WebRTC users. "For agents and telephony (where there is no browser frontend), use the LiveKit Cloud models."
Those models do different jobs:
| Tool | Removes | Helps with | Doesn't help with |
|---|---|---|---|
| Browser WebRTC AEC (client) | The agent's voice from the user's mic | Laptop speaker echo | PSTN callers |
| Background noise suppression (Krisp NC, ai-coustics QUAIL_L) | Non-speech noise | Fans, traffic, music beds | TV dialogue, side talk |
| Voice isolation (Krisp VIVA, VIVA telephony, ai-coustics QUAIL_VF_S/L) | Competing voices and noise | TV, side conversations, cafe chatter | Multi-speaker rooms where everyone counts |
| Krisp at the SIP trunk (`krisp_enabled`) | Noise, standard NC model only | Basic telephony noise | Anything voice isolation does |
Three rules come out of the docs. First, voice isolation, not noise suppression, is the fix for TV and side talk, because both are speech. Second, don't stack models. The docs warn that noise cancellation models "are trained on raw audio," so enabling them in both the frontend and the agent can produce unexpected results. Third, these enhanced models need LiveKit Cloud transport. Self-hosters have to bring their own filter or authenticate ai-coustics directly with a license key.
The docs show the effect on a noisy sample. Deepgram Nova-3 WER on the raw "gym membership" clip was 117.6%. That fell to 11.8% with Krisp VIVA and to 7.1% with ai-coustics QUAIL_VF_S. That's one clip, not a benchmark, but it shows how much of a "turn-taking bug" can be an audio bug.
Telephony vs WebRTC: why the same config behaves differently
| Factor | WebRTC (browser or app) | Inbound SIP | Outbound SIP |
|---|---|---|---|
| Client echo cancellation | Yes, recommended on | Depends on caller's handset | Depends on callee's handset |
| `aec_warmup_duration` default | 3.0 s | 3.0 s | `None` |
| Greeting interruptible from frame 1 | No | No | Yes |
| Audio bandwidth into the models | Wideband | Narrowband (G.711 carries nothing above 4 kHz) | Narrowband |
| Recommended NC per LiveKit tuning docs | Voice isolation | `BVCTelephony()` or VIVA telephony | Same |
| Speakerphone and car echo | Rare | Common | Common, plus voicemail greetings |
The adaptive model takes 16 kHz input, so narrowband phone audio reaches it upsampled with an empty top band. LiveKit says the model holds up across noise and languages. It doesn't publish a telephony-specific result. Measure on phone audio, because a config tuned in the browser console says little about a Twilio leg. Our SIP vs WebRTC comparison covers the transport differences.
Stop latency also differs by channel. The agent's `listening` state change is agent-side. The caller keeps hearing whatever audio is already buffered downstream. Measure stop latency from a caller-side recording, as the LiveKit testing guide explains for end-to-end latency.
Worked example: what one cough costs
Take the documented defaults in VAD mode: `min_duration` 0.5 s and `false_interruption_timeout` 2.0 s. Assume the Silero plugin's default end-of-speech silence of 0.55 s. A caller coughs for 0.6 s while the agent talks.
- t = 0.00 s: cough starts.
- t = 0.50 s: VAD speech duration reaches `min_duration`, so the agent pauses.
- t = 0.60 s: cough ends.
- t ≈ 1.15 s: VAD reports end of speech after 0.55 s of silence, and the 2.0 s timer arms.
- t ≈ 3.15 s: no transcript, so `agent_false_interruption` fires and speech resumes.
The caller hears 3.15 − 0.50 ≈ 2.65 seconds of dead air from one cough. Adaptive mode avoids the pause entirely when the model calls it a non-interruption, which is the main reason to make sure adaptive is actually running.

Now scale it. These are illustrative assumptions, not measurements. Say an agent talks for 2.5 minutes per call and callers produce 4 overlaps per agent-minute. That's 10 overlaps per call. At a 15% false-interruption rate, that's 1.5 stalls per call. At 5% it's 0.5. Shorter agent turns cut the overlap count, and better gating cuts the rate. The two multiply.
What the research says about the trade-off
There is no setting that lowers false interruptions without raising something else. Every gate that ignores more overlap also ignores some real barge-ins, or yields later.
Full-Duplex-Bench v1.5 measured this directly across five systems. Each one heard 200 synthetic interruptions, 99 backchannels, 100 side-speech clips and 100 background-speech clips. The paper scores the reply after each overlap as Respond, Resume, Uncertain or Unknown. It also scores stop latency (model stop minus user onset) and response latency (next model utterance minus user end).
| System | Interruption: Respond | Interruption stop latency | Backchannel: Resume | Background speech: Respond (lower is better) |
|---|---|---|---|---|
| GPT-4o Realtime (Dec 2024) | 0.78 | 0.23 s | 0.70 | 0.93 |
| Freeze-Omni | 0.72 | 1.42 s | 0.80 | 0.63 |
| Moshi | 0.50 | 1.16 s | 0.06 | 0.21 |
| Gemini 2.0 Flash Live | 0.33 | 2.20 s | 0.93 | 0.70 |
| Nova Sonic v1 | 0.24 | 2.25 s | 0.98 | 0.01 |
The authors describe "two divergent strategies": a responsive approach that yields fast, and a floor-holding approach that filters overlap. The fast yielder stopped in 0.23 s on real interruptions, and also answered 93% of background speech. The best filters resumed after 93-98% of backchannels, and took more than 2 seconds to yield to a real interruption. These are 2025 model versions, so read them as design evidence, not a current ranking.

LiveKit sits in the same trade-off space, with knobs instead of a fixed model. Raising `min_duration` or `min_words`, or widening the cooldown, moves you toward floor-holding. Lowering them moves you toward responsive. Adaptive mode is LiveKit's attempt to move the whole curve, and the launch numbers are worth reading carefully. "86% precision and 100% recall (at 500 ms overlap speech)" means that at that overlap length the model missed no real barge-ins, and about 14% of the stops it triggered weren't real interruptions. So the false-interruption resume path still matters even with adaptive mode on.
Our interruption detection evaluation guide covers general scoring definitions. The rest of this guide is LiveKit-specific.
Cause, symptom, fix
| Cause | Symptom you hear | How to confirm in logs | Fix |
|---|---|---|---|
| Adaptive mode not active (self-hosted `start`, STT without aligned transcripts, quota) | Every "mm-hm" over 0.5 s pauses the agent | No `overlapping_speech` events at all | Deploy to Cloud or run `dev`; check `stt.capabilities.aligned_transcript`; pin `mode` |
| Adaptive fell back to VAD mid-session | Good early in the call, jumpy later | `overlapping_speech` events stop while overlaps continue | Watch inference errors; budget the 40,000 free dev requests |
| Backchannel inside the 1 s start cooldown | Agent stops on "okay" right after it starts talking | Overlap onset within 1 s of `agent_state` → `speaking` | Lower `backchannel_boundary[0]` toward 0.5 s, test corrections |
| Coughs, sighs, clicks | Stall of about 2.5-3 s, then resume | `agent_false_interruption` with `resumed=True` | Adaptive mode; or raise `min_duration` to 0.6-0.8 s in VAD mode |
| TV or side talk | Agent abandons its answer and replies to nobody | Assistant item `interrupted=True` with an odd user transcript | Voice isolation (VIVA or QUAIL_VF), not noise suppression |
| Speakerphone echo | Agent stops itself, often in the first seconds | Overlaps start right as the agent starts speaking, with no caller speech in the recording | Telephony NC model; keep or lengthen `aec_warmup_duration` |
| `min_words` set above 0 | Caller says "No" over the agent and is ignored | Final transcript present, no user turn committed | Keep `min_words=0`; use adaptive mode for backchannels |
| False-interruption timer fires during slow speech | Agent resumes over a caller still talking; their words vanish | Resume while interim transcripts are still arriving | Don't drop `false_interruption_timeout` below 1.5 s; see issue #7198 |
| STT endpointing mode with old versions | Resume message after every turn, delayed replies | "resumed false interrupted speech" each turn | Upgrade; issue #4615 was closed in January 2026 |
| Greeting deaf window | Callers repeat themselves after talking over the greeting | `user_transcription_timeout` near call start | Shorten the greeting below 3 s, or set `aec_warmup_duration` deliberately |
Three of these rows deserve detail, because they're the ones teams don't find on their own.
`min_words` throws away answers, not just backchannels. The source applies the check twice. During overlap, `_interrupt_by_audio_activity` returns without pausing if the interim transcript has fewer than `min_words` words. When the user turn completes while the agent is still speaking, a transcript shorter than `min_words` cancels preemptive generation and returns without committing the turn. Our LiveKit testing guide flags it in its gotcha list too. With `min_words=2`, a caller who says "No" while the agent reads back a wrong date gets nothing. In a scheduling or payments flow, that's worse than a false interruption.
Adaptive mode can turn off without telling you. Unrecoverable inference errors make the session fall back to VAD interruption, and they aren't re-emitted as a session `error` event. The source treats a request with no answer after 700 ms as a non-retryable timeout. It also treats HTTP 429 as non-retryable "quota exceeded." The local-development allowance is 40,000 requests a month, and the client sends one request per 100 ms of overlap. Track the share of overlaps that produced an `overlapping_speech` event. If it drops, you're in VAD mode.
The docs example event doesn't exist in the event list. The turns overview shows a handler for `user_interruption_detected` with `ev.probability`. The session's `EventTypes` in current source has no such event. The documented, working event is `overlapping_speech` (`OverlappingSpeechEvent`), which carries `is_interruption`, `probability`, `detection_delay`, `overlap_started_at` and `num_requests`.
How to fix false interruptions in a LiveKit agent
Work in this order. Each step changes what the next step measures, so tuning thresholds before the audio is clean wastes effort.
1. Log first. Attach the event logger from the next section and record stereo audio (caller on one channel, agent on the other) for a baseline set of scripted calls. Without a baseline you can't tell a fix from noise.
2. Confirm the mode you think you have. Pin `interruption.mode` explicitly. Check that `overlapping_speech` events appear during overlaps. If they don't, check `stt.capabilities.aligned_transcript`, whether you're running `start` self-hosted without Cloud inference, and whether you've used up the dev quota.
3. Fix the audio path. Add voice isolation for single-caller lines, using the telephony variant for SIP participants. Remove any duplicate noise model in the frontend. Re-measure before touching thresholds. TV and echo problems often disappear here.
4. Decide your greeting policy. On inbound calls, either keep the greeting under 3 seconds or set `aec_warmup_duration` on purpose. On outbound calls, decide whether a callee's "hello?" should stop the opener, and set the warmup to match.
5. Tune the cooldown before the gates. In adaptive mode, most remaining backchannel stops come from the first second of each agent turn. Try `backchannel_boundary=(0.5, 1.0)` and re-run your correction scenarios ("no, wait") to confirm real barge-ins still get through.
6. Tune `min_duration` with a sweep, not a guess. Run 0.4, 0.5, 0.6 and 0.8 s against the same test set. In VAD mode it's your main control. In adaptive mode it mostly affects the cooldown window and fallback.
7. Leave `min_words` at 0 unless you've tested short answers spoken over agent speech and accept losing them.
8. Set the resume behavior deliberately. Keep `resume_false_interruption=True`. Lower `false_interruption_timeout` toward 1.5 s only if your test set shows no resumes over slow speakers. Never disable it in adaptive mode, because about 14% of model-triggered stops can still be false.
9. Re-run the full matrix and pick a point on the curve. Plot false-interruption rate against missed-barge-in rate for each config you tried. Pick the one that meets both thresholds below, not the one with the best single number.
10. Pin and guard. Commit the final values, add a config test that fails on drift, and re-run the overlap suite whenever you change STT, TTS voice, noise model or LiveKit version.
A starting config for a phone agent on LiveKit Cloud looks like this. The values are a reasonable first test point, not a recommendation for every agent:
# agent.py (illustrative; livekit-agents Python >= 1.5, LiveKit Cloud)
from livekit.agents import AgentSession, TurnHandlingOptions, inference, room_io
from livekit.plugins import noise_cancellation
session = AgentSession(
# ... stt, llm, tts, vad
aec_warmup_duration=3.0, # explicit: now also applies to outbound SIP (auto-default is None there)
turn_handling=TurnHandlingOptions(
turn_detection=inference.TurnDetector(),
interruption={
"mode": "adaptive", # pinned; the session falls back to "vad" if unavailable
"min_duration": 0.5, # gates VAD path: cooldown window and fallback
"min_words": 0, # >0 discards short answers spoken over the agent
"false_interruption_timeout": 2.0,
"resume_false_interruption": True,
"backchannel_boundary": (0.5, 1.0), # Python only: start, end cooldown
},
),
)
await session.start(
# ... room, agent
room_options=room_io.RoomOptions(
audio_input=room_io.AudioInputOptions(
noise_cancellation=noise_cancellation.BVCTelephony(), # SIP callers
),
),
)Measuring false interruptions: definitions and formulas
You need four numbers. Each needs a labeled stimulus, meaning you know what you injected and when.
False-interruption rate (FIR). For non-floor-taking stimuli (backchannels, coughs, TV, side speech, echo) played during agent speech:
FIR_any = stimuli that stopped agent audio ÷ non-floor-taking stimuli
FIR_hard = stimuli that stopped agent audio and did not resume ÷ non-floor-taking stimuli
Report both. A resumed pause costs a stall. A hard false interruption costs the answer and the conversation state.
Missed-barge-in rate (MBR). For real barge-ins, such as corrections and new requests:
MBR = barge-ins where agent audio didn't stop within 1.5 s of onset ÷ real barge-ins
The 1.5 s window is a choice, not a standard. Pick one and keep it fixed across runs.
Stop latency. Agent audio stop minus caller onset, from the caller-side recording, reported as p50 and p95. This is the `t_stop` metric from Full-Duplex-Bench v1.5. Agent-side `agent_state_changed` timestamps understate it by whatever is buffered downstream.
Resume gap. For resumed false interruptions, agent audio restart minus stimulus end. With defaults, expect roughly end-of-speech silence plus 2.0 s.
Two supporting numbers keep the main ones honest:
- Adaptive coverage = overlaps with an `overlapping_speech` event ÷ overlaps detected by VAD. Below about 95%, you're partly in VAD mode.
- Swallowed-turn rate = short real answers spoken over the agent's last second that never became a committed user turn ÷ such answers. This catches the end-boundary and `min_words` problems.
How many stimuli you need
To estimate a false-interruption rate near 10% within ±3 points at 95% confidence: n = 1.96² × 0.10 × 0.90 ÷ 0.03² = 3.8416 × 0.09 ÷ 0.0009 ≈ 384 stimuli.
To detect a config change that moves FIR from 15% to 8% at 5% significance and 80% power: n per arm = (1.96 + 0.84)² × (0.15 × 0.85 + 0.08 × 0.92) ÷ 0.07² = 7.84 × 0.2011 ÷ 0.0049 ≈ 322 stimuli per config.
For missed barge-ins, where you hope to see zero, use the rule of three. Zero misses in n trials puts the 95% upper bound near 3/n. Zero misses in 60 barge-ins means the true rate could still be 5%. Our guide to A/B testing voice agents covers the same math for whole-call outcomes.
Logging LiveKit interruption events
This logger writes one JSON line per relevant event, keyed to wall-clock time. Every event and field name comes from the events reference or the `events.py` source.
# interruption_log.py (illustrative; livekit-agents Python >= 1.5)
import json
import time
from pathlib import Path
from livekit.agents import AgentSession
def attach_interruption_logger(session: AgentSession, call_id: str, out_dir: str = "logs"):
path = Path(out_dir) / f"{call_id}.jsonl"
path.parent.mkdir(parents=True, exist_ok=True)
fh = path.open("a")
def write(kind: str, at: float | None = None, **fields) -> None:
rec = {"call_id": call_id, "kind": kind, "at": at or time.time(), **fields}
fh.write(json.dumps(rec, default=str) + "\n")
fh.flush()
@session.on("agent_state_changed")
def _agent_state(ev):
write("agent_state", ev.created_at, old=ev.old_state, new=ev.new_state)
@session.on("user_state_changed")
def _user_state(ev):
write("user_state", ev.created_at, old=ev.old_state, new=ev.new_state)
@session.on("overlapping_speech") # adaptive mode only
def _overlap(ev):
write("overlap", ev.detected_at, is_interruption=ev.is_interruption,
probability=ev.probability, detection_delay=ev.detection_delay,
overlap_started_at=ev.overlap_started_at, num_requests=ev.num_requests)
@session.on("agent_false_interruption")
def _false_interruption(ev):
write("false_interruption", ev.created_at, resumed=ev.resumed)
@session.on("user_input_transcribed")
def _transcript(ev):
if ev.is_final:
write("user_final", ev.created_at, text=ev.transcript)
@session.on("user_transcription_timeout")
def _tx_timeout(ev):
write("transcription_timeout", ev.created_at, speech_duration=ev.speech_duration)
@session.on("conversation_item_added")
def _item(ev):
item = ev.item
if getattr(item, "role", None) == "assistant":
write("assistant_item", ev.created_at, interrupted=item.interrupted,
text=item.text_content)
@session.on("close")
def _close(ev):
write("close", reason=ev.reason)
fh.close()Call `attach_interruption_logger(session, ctx.room.name)` before `session.start()`. Note that `user_transcription_timeout` only fires if you set `transcription_timeout` on `AgentSession`. It's off by default.
Scoring against a stimulus schedule
Your test harness knows when it played each stimulus. Write that to a CSV with `call_id`, `stimulus_id`, `kind` and `onset` and `offset` in epoch seconds, using the same clock as the agent or an NTP-synced one. Then score:
# score_interruptions.py (illustrative)
import csv
import json
import math
import statistics
from collections import defaultdict
STOP_WINDOW = 1.5 # s after onset to count a stop
RESUME_WINDOW = 6.0 # s after onset to look for a resume
NON_FLOOR = {"backchannel", "cough", "tv", "side_speech", "echo"}
def wilson(k: int, n: int, z: float = 1.96) -> tuple[float, float]:
if n == 0:
return (0.0, 1.0)
p = k / n
d = 1 + z * z / n
c = p + z * z / (2 * n)
h = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n))
return ((c - h) / d, (c + h) / d)
def load_events(paths):
by_call = defaultdict(list)
for p in paths:
for line in open(p):
e = json.loads(line)
by_call[e["call_id"]].append(e)
for evs in by_call.values():
evs.sort(key=lambda e: e["at"])
return by_call
def score(stimuli_csv: str, event_files: list[str]) -> dict:
events = load_events(event_files)
rows = []
for s in csv.DictReader(open(stimuli_csv)):
onset, evs = float(s["onset"]), events.get(s["call_id"], [])
stop = next((e for e in evs if e["kind"] == "agent_state"
and e["old"] == "speaking" and onset <= e["at"] <= onset + STOP_WINDOW), None)
resumed = any(e["kind"] == "false_interruption" and e["resumed"]
and onset <= e["at"] <= onset + RESUME_WINDOW for e in evs)
rows.append({"kind": s["kind"], "stopped": stop is not None, "resumed": resumed,
"stop_latency": stop["at"] - onset if stop else None})
nf = [r for r in rows if r["kind"] in NON_FLOOR]
bi = [r for r in rows if r["kind"] == "barge_in"]
k_any = sum(r["stopped"] for r in nf)
k_hard = sum(r["stopped"] and not r["resumed"] for r in nf)
k_miss = sum(not r["stopped"] for r in bi)
lat = sorted(r["stop_latency"] for r in bi if r["stop_latency"] is not None)
return {
"fir_any": (k_any / max(len(nf), 1), wilson(k_any, len(nf))),
"fir_hard": (k_hard / max(len(nf), 1), wilson(k_hard, len(nf))),
"mbr": (k_miss / max(len(bi), 1), wilson(k_miss, len(bi))),
"stop_p50_agent_side": statistics.median(lat) if lat else None,
"stop_p95_agent_side": lat[int(0.95 * (len(lat) - 1))] if lat else None,
}Two cautions. First, only place stimuli where the agent has at least 1.5 s of speech left. Otherwise a natural end of speech looks like a stop. Second, these stop latencies are agent-side. For the caller's view, run a VAD over the agent channel of the stereo recording and take the first frame of silence after onset. Our guide on what to log on every voice agent call covers recording layout.
A test set for LiveKit false interruptions
This is the matrix we'd use as a starting point. It's named so you can refer to it in tickets: the LFI-9 overlap set. It has nine stimulus types, each placed during agent speech:
| ID | Stimulus | Floor-taking? | Placement |
|---|---|---|---|
| S1 | Short backchannel: "mm-hm," "uh-huh" (under 0.5 s) | No | About 700 ms after an agent phrase boundary, more than 1 s into the turn |
| S2 | Long backchannel: "yeah, okay, sure," "right, right" (over 0.5 s) | No | Same |
| S3 | Backchannel in the start cooldown | No | Within the first 1 s of an agent turn |
| S4 | Cough, sneeze, throat clear | No | Mid-turn |
| S5 | Side speech ("hang on, honey"), using the FDB v1.5 far-field recipe | No | Mid-turn |
| S6 | TV or background speech, using the FDB v1.5 background recipe | No | Mid-turn, 3-6 s long |
| S7 | Real barge-in correction ("no, wait, Thursday not Tuesday") | Yes | Mid-turn and in the first 1 s |
| S8 | Short real answer ("no") in the agent's last second | Yes | Over the end of a yes/no question |
| S9 | Caller talks over the greeting | Yes | In the first 3 s of the call |
Run each stimulus under four conditions:
- C1 clean.
- C2 noise bed at 10 dB SNR (cafe or street). To mix at a target SNR, scale the noise by g = RMS_speech ÷ (RMS_noise × 10^(SNR/20)).
- C3 noise bed at 5 dB SNR.
- C4 speakerphone echo. Re-inject the agent's own TTS into the caller channel. As an illustrative starting point, try -15 dB at 80 ms delay, then vary both.
Run all of it over the channel you ship, PSTN through your SIP trunk, not just the browser console. Thirty trials per cell gives 9 × 4 × 30 = 1,080 stimuli. At five or six stimuli per scripted call, that's about 200 calls, which fits a nightly run.
Pass thresholds
These are starting thresholds for a support or scheduling agent. They're our recommendation, not an industry standard, so adjust them for your callers:
| Metric | Cells | Pass |
|---|---|---|
| FIR_hard | S1, S2 after the cooldown, C1-C3 | ≤ 5% |
| FIR_any | S1, C1 | ≤ 10% |
| FIR_any | S3 (cooldown) | Record it; expect it to be higher by design |
| FIR_hard | S4 coughs, all conditions | ≤ 2% |
| FIR_hard | S5, S6 side and TV speech, C1-C2 | ≤ 10% |
| Self-interruptions | C4 echo, no caller stimulus | 0 per call |
| MBR | S7, C1 | 0 misses in at least 60 (upper bound about 5%) |
| MBR | S7, C3 | ≤ 5% |
| Stop latency, caller-side | S7, C1 | p95 ≤ 800 ms |
| Swallowed turns | S8 | ≤ 5% |
| Greeting capture | S9, inbound | 100% of caller words transcribed, or a deliberate policy documented |
| Resume gap | Resumed S1, S2, S4 | p95 ≤ 3.0 s at default timeout |
| Context integrity | S7 followed by "what did you just tell me?" | Agent repeats only what was heard: 100% |
Read variance before you read means. Run each cell twice on different days. If the same config's FIR moves more than the Wilson interval suggests, something in the environment is changing: a different region, adaptive fallback, or a provider update. Fix that before comparing configs. Our VAD misfire testing guide and background noise STT guide cover building the noise beds.
Known issues worth checking against your version
Three public issues map directly onto the failure modes above.
- #4615, closed January 2026. With Deepgram STTv2 and `turn_detection="stt"` on livekit-agents 1.3.12, every user turn produced a phantom "resumed false interrupted speech" and delayed playback by the 2.0 s timeout. Current source has an explicit guard for the case where the STT end-of-speech arrives before VAD's. If you're on an older version with STT endpointing, upgrade before tuning.
- #7198, reported on 1.7.1 with Deepgram Flux. Interim transcripts don't clear the false-interruption timer, so a caller who interrupts and then pauses mid-sentence can have the agent resume over them. Their interim words are lost. Maintainers said interims shouldn't clear the timer, since noise also produces them. The thread discusses re-arming with a cap instead. Your S7 tests with mid-sentence pauses will catch this.
- #5038, reported on 1.4.4. With `resume_false_interruption=True`, an agent reply paused before its first audio frame played, then truly interrupted, could vanish from conversation history. The user heard part of it, but `chat_ctx` didn't record it. The reporter's workaround was disabling resume. The context-integrity probe in the threshold table is how you detect this class of bug.
When to bring in independent testing
Everything above can be done in-house, and you should own the logger and a nightly run. Independent evaluation helps in three places: before launch, when you change a component, and when the team that tuned the config is also the team grading it.
A pre-launch audit runs the overlap matrix over real carriers and real handsets, including speakerphones and cars, which a laptop console can't reproduce. A vendor or config bake-off runs two STT providers, or adaptive vs VAD, against identical stimuli, so the comparison isn't biased by whoever scripted the test. Regression and production scoring keep FIR and missed barge-ins on a dashboard after each LiveKit or provider upgrade. Evalgent does this as a third party, so the numbers come from someone who didn't write the config. See what independent voice AI evaluation covers, and track the production side with interruption rate.
Frequently asked questions
What is a false interruption in LiveKit?
A false interruption happens when LiveKit pauses the agent for user audio that wasn't a real attempt to take the turn, such as a backchannel, cough, TV or echo. If no transcript arrives within `false_interruption_timeout` (2.0 s by default) after VAD end of speech, the session emits `agent_false_interruption` and resumes. If a transcript does arrive, the false interruption becomes a hard one.
What does LiveKit adaptive interruption handling do?
It sends overlapping user audio to a barge-in model on LiveKit Cloud. The model streams 16 kHz windows every 100 ms and decides whether the user is interrupting or backchanneling. It needs a VAD plus an STT with aligned transcripts, or a realtime model with server turn detection off. It's on by default on LiveKit Cloud and in `dev`.
Why does my self-hosted LiveKit agent still stop on "mm-hm"?
It's probably running in VAD mode. Adaptive interruption runs on LiveKit Cloud inference, so a self-hosted agent started with `start` defaults to VAD-based interruption. Any "mm-hm" or "yeah" longer than `min_duration` (0.5 s) will pause it. Check for `overlapping_speech` events. If there are none, you're in VAD mode and `min_duration` is your main control.
What should I set min_duration to?
Start at the default 0.5 s and sweep 0.4, 0.6 and 0.8 s against a labeled test set. Higher values ignore more coughs and long backchannels but raise stop latency and the risk of missing short corrections. In adaptive mode, `min_duration` mostly applies during the first-second cooldown and after fallback, so tune `backchannel_boundary` there first.
Should I raise min_words to stop backchannel interruptions?
Usually not. The current source uses `min_words` twice. It blocks the pause, and it also drops the user turn entirely if the final transcript is shorter than `min_words` while the agent is speaking. With `min_words=2`, a caller saying "No" over a wrong readback is ignored. Use adaptive mode for backchannels and keep `min_words` at 0.
Why does my agent ignore callers who talk over the greeting?
On inbound calls and WebRTC sessions, `aec_warmup_duration` defaults to 3.0 s. During that window on the first agent turn, interruptions are disabled and STT receives silence instead of caller audio. Outbound SIP calls default to no warmup. Keep greetings under 3 s, or set `aec_warmup_duration` explicitly after testing echo.
Does noise cancellation fix false interruptions from TV?
Only voice isolation does. TV dialogue and side conversations are speech, so background noise suppression keeps them. Use Krisp VIVA (the telephony variant for SIP) or ai-coustics Voice Focus. Don't stack noise models in both the frontend and the agent. LiveKit's docs warn that the models expect raw audio.
How do I measure false interruption rate on a LiveKit agent?
Play labeled stimuli during agent speech and log `agent_state_changed`, `agent_false_interruption`, `overlapping_speech` and assistant items with `interrupted`. FIR is stimuli that stopped agent audio divided by non-floor-taking stimuli. Report resumed and hard stops separately, and take stop latency from caller-side recordings. About 384 stimuli give ±3 points at a 10% rate.
The bottom line
Most LiveKit false interruptions come from a mode you didn't know you were in, an audio path nobody cleaned, or a default that behaves differently on phone calls, not from bad thresholds. Log the interruption events, run a labeled overlap set over the channel you actually ship, and tune only when you can show the false-interruption rate dropping while missed barge-ins and stop latency stay in budget.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more