Evalgent
Back to Blog
Voice AI Testing

Pipecat Interruption Strategy and Smart Turn v3: Settings That Work on Phone Audio

Deepesh Jayal
22 min read
Pipecat Interruption Strategy and Smart Turn v3: Settings That Work on Phone Audio
On this page

Most turn-taking problems on a Pipecat phone agent come from three settings that interact: the Silero VAD timers, the start strategy that decides when the caller has interrupted, and the stop strategy that decides when the caller is done. The official docs describe each class. They do not tell you how the three behave together on 8 kHz μ-law audio, where callers pause mid-sentence, read card numbers in chunks, and say "mm-hm" while the bot talks.

This guide works from the `pipecat-ai` 1.12.0 source on PyPI, the Pipecat changelog, the Smart Turn model card and benchmark files, and the turn-taking research those models build on. Every class and parameter below exists in 1.12.0. Numbers are either sourced or labeled as worked examples.

This page does not repeat our Pipecat voice agent testing guide, which covers the `min_volume` loudness gate, per-provider STT latency constants, and Pipecat Evals. It goes one level deeper on interruption and end-of-turn behavior.

8 s
Max audio Smart Turn v3 scores per decision (smart-turn README)
3.75%
English false "complete" rate, Smart Turn v3.2 CPU benchmark
59.6%
Within-speaker pauses longer than 500 ms, Swedish Map Task (Heldner and Edlund 2010)
65 ms
Smart Turn v3 inference on a Pipecat Cloud 1x instance (smart-turn README)

What happened to allow_interruptions and MinWordsInterruptionStrategy

If you search for a Pipecat interruption strategy, most results still show `PipelineParams(allow_interruptions=True, interruption_strategies=[MinWordsInterruptionStrategy(min_words=3)])`. That code no longer does what it says. The `MinWordsInterruptionStrategy` import fails outright. Worse, `PipelineParams` and `TransportParams` are pydantic models with the default "ignore extra fields" behavior, so a leftover `allow_interruptions=False` or `vad_analyzer=` argument is silently dropped with no error. Pipecat 0.0.99 (January 13, 2026) deprecated it, and 1.0.0 (April 14, 2026) removed it. Per the changelog, interruptions are now always allowed by default, and you control them through the user aggregator.

Old API (0.0.98 and earlier)Pipecat 1.x replacementWhere it lives
`PipelineParams.allow_interruptions=False`Start strategies with `enable_interruptions=False`, or a mute strategy`LLMUserAggregatorParams.user_turn_strategies` / `user_mute_strategies`
`interruption_strategies=[MinWordsInterruptionStrategy(min_words=3)]``MinWordsUserTurnStartStrategy(min_words=3)``UserTurnStrategies(start=[...])`
`TransportParams(vad_analyzer=...)``LLMUserAggregatorParams(vad_analyzer=...)`User aggregator
`TransportParams(turn_analyzer=...)``TurnAnalyzerUserTurnStopStrategy(turn_analyzer=...)``UserTurnStrategies(stop=[...])`
`STTMuteFilter``FunctionCallUserMuteStrategy`, `MuteUntilFirstBotCompleteUserMuteStrategy`, and others`user_mute_strategies`
`UserIdleProcessor``user_idle_timeout` plus the `on_user_turn_idle` eventUser aggregator

The changelog's own example for turning interruptions off while still getting user turns is a start list of `[TranscriptionUserTurnStartStrategy(enable_interruptions=False)]`. That works, but on a phone line it means the bot talks over every caller who says "wait." Most phone agents want interruptions on, filtered.

How a user turn moves through Pipecat 1.12

Everything below happens inside `LLMUserAggregator`, the user half of `LLMContextAggregatorPair`. Since 1.0 it owns VAD, turn start, turn stop, idle detection and muting. Knowing the order of operations explains most of the bugs you will see on calls.

Pipecat 1.12 user turn flow: Silero VAD windows feed start strategies that broadcast an InterruptionFrame, then Smart Turn runs at each VAD stop before release

VAD: 32 ms windows, counted in whole windows

`SileroVADAnalyzer` accepts only 8,000 or 16,000 Hz. It scores 256-sample windows at 8 kHz or 512-sample windows at 16 kHz, so each window is 32 ms either way. A window counts as speech only when Silero's confidence is at least `confidence` and the smoothed loudness is at least `min_volume`.

`start_secs` and `stop_secs` are converted to window counts with `round(secs / 0.032)`. The 1.12.0 defaults are `confidence=0.7`, `start_secs=0.2`, `stop_secs=0.2`, `min_volume=0.6`. That gives 6 windows, or 192 ms, to declare speech and 6 windows to declare silence. The `stop_secs` default dropped from 0.8 to 0.2 in 0.0.102 (February 2026), because Smart Turn now makes the end-of-turn call and VAD only needs to notice the pause.

On Twilio, `TwilioFrameSerializer` decodes 8 kHz μ-law and resamples it to the pipeline rate, which is `PipelineParams.audio_in_sample_rate=16000` by default. So in the default setup, Silero and Smart Turn both see band-limited audio that has been upsampled. Nothing above 4 kHz survives the phone network. Upsampling changes the frame size, not the content.

Start strategies decide what counts as an interruption

The defaults are `[VADUserTurnStartStrategy(), TranscriptionUserTurnStartStrategy()]`. The controller runs the start list in order, and the first strategy that returns `STOP` opens the turn. When a turn opens with `enable_interruptions=True`, the aggregator does three things: it broadcasts `UserStartedSpeakingFrame`, tells the idle controller the user is active, and calls `broadcast_interruption()`.

With the VAD strategy first, a caller interrupts after about 192 ms of detected speech, before any word is transcribed. The transcription strategy is a fallback for soft speech the VAD misses.

Stop strategies and where Smart Turn runs

The default stop list is `[TurnAnalyzerUserTurnStopStrategy(turn_analyzer=LocalSmartTurnAnalyzerV3())]`. Smart Turn does not run on every audio frame. It runs once each time VAD reports `VADUserStoppedSpeakingFrame`, which means at every pause of about 200 ms or more.

1. VAD reports a stop. The strategy computes an absolute deadline: speech end plus the STT service's `ttfs_p99_latency`.

2. Smart Turn scores the buffered turn audio. Probability above 0.5 means complete.

3. If the verdict is complete, the turn releases as soon as a finalized transcript arrives. Deepgram's Pipecat service sends a Finalize message on VAD stop and marks the reply as finalized. If no finalized transcript arrives, the turn releases at the deadline.

4. If the verdict is incomplete, the turn stays open. Smart Turn's own `SmartTurnParams.stop_secs` (default 3 seconds of silence) forces it closed if the caller says nothing more.

5. If nothing closes the turn, `user_turn_stop_timeout` (default 5.0 seconds) ends it and fires `on_user_turn_stop_timeout`.

The docs say this plainly, and it matters for latency: VAD `stop_secs` and Smart Turn inference fall inside the STT deadline rather than adding to it. We use that in the latency budget below.

What an interruption does to frames and the LLM context

`InterruptionFrame` is a system frame, so processors handle it immediately instead of queuing it behind audio. The LLM stops generating, TTS stops, and the output transport drops its buffered audio. On Twilio, the serializer converts the interruption into a `{"event": "clear"}` message so Twilio flushes the audio it has buffered on its side.

The assistant aggregator sits after `transport.output()`. It receives the text frames the TTS emits as audio plays out. On interruption it closes the assistant turn with `interrupted=True` and commits the text aggregated so far to the context. With word-timestamped TTS, that is close to what the caller actually heard. The testing guide covers the cases where context drifts from what the caller heard.

One more 1.x behavior worth knowing: `empty_user_turn`. If a turn interrupted the bot but produced no transcript (a cough, a door), the default `EmptyUserTurnConfig` runs the LLM once with a prompt asking the caller to repeat. Without it, the bot would stop mid-sentence and stay silent.

Smart Turn v3.2: what the model actually sees

The smart-turn repository describes the model as Whisper Tiny with a linear classifier, about 8M parameters. It ships as an 8 MB int8 CPU build and a 32 MB fp32 GPU build, and supports 23 languages. Pipecat 1.12.0 bundles `smart-turn-v3.2-cpu.onnx`.

The source of `LocalSmartTurnAnalyzerV3` shows the exact input pipeline:

  • Audio buffered from 500 ms before VAD speech start (`pre_speech_ms`), plus the VAD's own `start_secs`, so the first syllable is not cut off.
  • Resampled to 16 kHz with soxr if the pipeline runs at another rate.
  • Truncated to the last 8 seconds, or zero-padded at the front if shorter.
  • Converted to Whisper-style log-mel features, scored, thresholded at 0.5.
  • Run on one dedicated thread per analyzer, with `cpu_count=1` intra-op threads by default.

Two consequences follow. A caller who talks for 12 seconds before pausing is judged only on the final 8. And because the model uses audio only, it reads prosody (falling pitch, final lengthening) rather than words. Pitch survives the phone band well, which is good news for 8 kHz lines.

What the public benchmark covers, and what it leaves out

The v3.2 CPU benchmark report is unusually transparent. Read it before you trust the model on your calls. The model's positive class is "complete," so the report's false-positive rate is the share of unfinished turns judged finished. Those are your premature cutoffs.

SliceSamplesAccuracyFalse "complete" rateFalse "incomplete" rate
Overall31,52792.63%4.73%2.64%
English7,82094.26%3.75%1.99%
Spanish1,78389.57%6.95%3.48%
Mid-utterance fillers (orpheus_midfiller_1)14087.14%7.14%5.71%
Human conversation (human_convcollector_1)9086.67%8.89%4.44%

Three things stand out for a US phone agent:

1. Most of the test set is synthetic. The two `chirp3` datasets hold 24,682 of 31,527 samples, about 78%. The dataset names indicate TTS-generated speech. The human conversation slice is 90 samples, and its premature rate is more than double the English average.

2. No dataset is identified as telephony. The v3.2 release added cafe and office noise, per Daily's release notes. Nothing in the report suggests 8 kHz μ-law audio. Treat phone accuracy as unknown until you measure it.

3. Spanish is noticeably weaker than English. If 20% of your callers speak Spanish, expect their premature-cutoff rate to be close to double.

The repository's roadmap also lists "text conditioning of the model, to support modes like credit card, telephone number, and address entry" as a future goal. In other words, today's model has no idea the caller is reading digits. Your configuration has to supply that context.

Why silence timers alone cannot fix this: what the research says

Pauses inside a turn are often longer than gaps between turns. Heldner and Edlund (Journal of Phonetics, 2010) measured silences in Swedish, Scottish English and Dutch conversation corpora. In the two Map Task corpora, a 500 ms silence threshold captured 51.1% and 47.5% of between-speaker gaps, but also 59.6% and 56.0% of within-speaker pauses. A 1,000 ms threshold still captured 27.7% to 31.0% of pauses. Because there were more pauses than gaps, a pure silence timer fires on more mid-turn pauses than true turn ends. The most common between-speaker interval in all three corpora was a gap of about 200 ms. That is why `SpeechTimeoutUserTurnStopStrategy` forces a bad trade: shorten the timer and you cut people off, lengthen it and every reply feels slow.

Expected gap length varies by language. Stivers et al. (PNAS, 2009) found the same overlap-avoiding, silence-minimizing pattern in 10 languages, with average gaps varying within about 250 ms of the cross-language mean. One timing setting tuned on English callers will feel slightly off for others. Combined with the Spanish benchmark row above, this is a reason to report premature-turn rate by language.

Phone numbers come in chunks, with pauses between chunks. Baumann and Trouvain (Eurospeech 2001) recorded speakers dictating German phone numbers to someone writing them down. Speakers grouped digits into chunks of two or three, and the same number was spoken in four different groupings on average. Non-final groups ended with a rising or level pitch, and only the last group fell. Their synthesis model used 450 ms pauses at group boundaries. US callers group a 10-digit number as 3-3-4, but the mechanism is the same: every group boundary is a VAD stop, and every VAD stop is a Smart Turn decision.

Noise type matters more than noise level. Russell and Harte (2025) tested Voice Activity Projection models, the turn-taking family from Ekstedt and Skantze (Interspeech 2022), with added noise. Competing speech and music at 10 dB SNR pushed hold-versus-shift accuracy to chance, while babble hurt far less. Smart Turn is a different model, but it is also audio-only. Test with a TV in the background, not only cafe noise.

For a model-agnostic view of how to score these behaviors, see our posts on turn-taking evaluation for voice models and pause handling. Full-Duplex-Bench uses the same split: pause handling, backchanneling, smooth turn-taking and user interruption.

The compounding problem: one turn, many chances to cut in

Smart Turn's benchmark error is per decision. Your caller's experience is per turn. Because Smart Turn runs at every pause of 200 ms or more, a turn with several internal pauses gives it several chances to call the turn finished too early.

If each decision on unfinished audio has an independent false "complete" probability `f`, and a turn contains `k` internal pauses, then:

`P(at least one premature cutoff) = 1 - (1 - f)^k`

Using the benchmark's English rate of 3.75% and Spanish rate of 6.95% as worked-example values (your phone audio may be better or worse):

Internal pauses `k`Typical utteranceEnglish, f = 3.75%Spanish, f = 6.95%
1"I need to, um, reschedule"3.7%7.0%
310-digit phone number in 3-3-4 groups, with a lead-in10.8%19.4%
6Street address plus apartment and ZIP20.5%35.1%
8Card number read in groups of four26.3%43.8%
Chart of the probability of at least one premature turn cutoff versus the number of internal pauses, for English and Spanish per-decision error rates from the Smart Turn v3.2 benchmark

Independence is an assumption. Errors on one caller's pauses are probably correlated, which could push the real numbers either way. The direction is not in doubt, though: long, structured utterances are where Smart Turn's small per-decision error becomes a large per-turn error. That is exactly the content of scheduling, intake, payments and address capture. It is also why aggregate dashboards look fine while number capture keeps failing. Our post on STT entity accuracy covers the transcription half of the same problem.

The MinWords trap: why "no" stops working

`MinWordsUserTurnStartStrategy(min_words=3)` is the standard advice for keeping "okay" and "mm-hm" from cutting off the bot. Reading the 1.12.0 source shows four behaviors the docs only partly state:

1. The threshold applies only while the bot speaks. When the bot is silent, one word opens a turn. That part is documented.

2. Words are counted per transcription frame, not across frames. A 0.0.100 fix (PR #3462) stopped the strategy from accumulating text across frames. "No... wait... stop," spoken with pauses long enough for the STT to finalize each word, never reaches three words.

3. Below-threshold words are discarded. When the count falls short, the strategy calls `trigger_reset_aggregation()`, and the aggregator resets. The caller's "no" never reaches the LLM context.

4. It replaces VAD as the barge-in trigger. The documented example uses `start=[MinWordsUserTurnStartStrategy(min_words=3)]` alone. Barge-in latency becomes STT interim latency plus the time to say three words. Assuming 2.5 words per second, that is about 1.2 seconds of the bot talking over the caller. With VAD start, it is about 0.2 seconds.

The result is a bot that ignores "no," "stop," and "wait," the three words a caller is most likely to use to interrupt. For collections and healthcare agents, an ignored "stop" is a compliance problem, not just a UX one.

There is also a version gate. Before 1.7.0, pairing a transcript-started turn (MinWords) with Smart Turn could end the user turn on every finalized transcript fragment, so the bot answered partial phrases (PR #5159). If you use MinWords with Smart Turn, run 1.7.0 or later.

A keyword escape for MinWords (illustrative)

This start strategy keeps the multi-word rule for backchannels and lets a short list of words interrupt immediately. It uses only public hooks from `BaseUserTurnStartStrategy`.

# Illustrative: verified against pipecat-ai 1.12.0 class and method names.
import re

from pipecat.frames.frames import (
    BotStartedSpeakingFrame,
    BotStoppedSpeakingFrame,
    InterimTranscriptionFrame,
    TranscriptionFrame,
)
from pipecat.turns.types import ProcessFrameResult
from pipecat.turns.user_start import BaseUserTurnStartStrategy

ESCAPE_WORDS = {"stop", "no", "wait", "hold", "agent", "representative", "operator", "human"}


class MinWordsWithEscapeStartStrategy(BaseUserTurnStartStrategy):
    """Require min_words to interrupt, except for escape words."""

    def __init__(self, *, min_words: int = 3, escape_words=ESCAPE_WORDS,
                 use_interim: bool = True, **kwargs):
        super().__init__(**kwargs)
        self._min_words = min_words
        self._escape = {w.lower() for w in escape_words}
        self._use_interim = use_interim
        self._bot_speaking = False

    async def handle_user_turn_started(self):
        self._bot_speaking = False

    async def process_frame(self, frame) -> ProcessFrameResult:
        if isinstance(frame, BotStartedSpeakingFrame):
            self._bot_speaking = True
        elif isinstance(frame, BotStoppedSpeakingFrame):
            self._bot_speaking = False
        elif isinstance(frame, TranscriptionFrame) or (
            self._use_interim and isinstance(frame, InterimTranscriptionFrame)
        ):
            words = re.findall(r"[a-z']+", frame.text.lower())
            if not words:
                return ProcessFrameResult.CONTINUE
            if (not self._bot_speaking
                    or len(words) >= self._min_words
                    or self._escape.intersection(words)):
                await self.trigger_user_turn_started()
                return ProcessFrameResult.STOP
            await self.trigger_reset_aggregation()
        return ProcessFrameResult.CONTINUE

Keep "okay," "yeah," and "uh-huh" out of the escape list, since they are the backchannels you are filtering. If your vendor budget allows it, `KrispVivaIPUserTurnStartStrategy` is the model-based alternative: Krisp's interruption-prediction model scores each VAD-detected segment as a real interruption or a backchannel.

Latency budget: what stop_secs and Smart Turn actually cost

This worked example traces the time from the caller's last syllable to the first bot audio leaving Pipecat. Assumptions are marked. The Deepgram figure is Pipecat's built-in `DEEPGRAM_TTFS_P99 = 0.35` s, measured with `stop_secs=0.2`. Smart Turn inference uses the README's 65 ms on a Pipecat Cloud 1x instance.

StageA: SpeechTimeout 0.6 sB: Smart Turn defaultC: Smart Turn, capture modeD: Smart Turn false "incomplete"
VAD stop window200 ms200 ms700 ms200 ms
End-of-turn decision+600 ms timer65 ms inference, hidden inside STT wait65 ms inference3,000 ms fallback
Turn released at (from speech end)about 800 msabout 350 ms (p99 STT deadline)about 765 ms or laterabout 3,200 ms
LLM time to first token (assumed)350 ms350 ms350 ms350 ms
TTS time to first byte (assumed)150 ms150 ms150 ms150 ms
Total before carrier legabout 1,300 msabout 850 msabout 1,265 msabout 3,700 ms
Latency waterfall comparing SpeechTimeout, Smart Turn default, Smart Turn in number-capture mode, and a Smart Turn false incomplete, broken into VAD, decision, LLM and TTS stages in milliseconds

Read three things from this table:

  • Smart Turn is nearly free on the happy path. Its 65 ms runs while the STT is still finalizing, so the release time is set by the transcript, not the model. Compared with a 0.6 s speech timeout, configuration B saves about 450 ms per turn.
  • Its expensive failure is the slow one. A false "incomplete" (1.99% of English decisions in the benchmark) costs the full 3-second fallback. Callers hear dead air and often say "hello?", which starts a new turn and confuses the context.
  • Capture mode costs about 400 to 500 ms, but only on the turns that use it. That is the trade we recommend for numbers and addresses below.

Add the telephony leg on top. Pipecat cannot measure it. Our LiveKit vs Pipecat latency breakdown shows how to measure it from the caller's side.

Find out how your Pipecat turn settings behave on real phone calls
Evalgent runs scripted callers over your actual Twilio or Telnyx number and reports premature cutoffs, missed barge-ins and response latency per scenario, so you can change one setting and see what it did.
Book a demo

Settings for 8 kHz phone audio

The table gives starting points to test, not measured optima. Change one row at a time and rerun the test set in the next section.

Setting1.12.0 defaultPhone starting pointWhy
`VADParams.confidence`0.70.7Raising it barely filters background voices, since Silero scores any human voice highly. It does lose soft callers.
`VADParams.start_secs`0.20.2, or 0.3 on noisy lines0.3 rounds to 9 windows (288 ms). It filters short clicks and line noise but adds 96 ms to barge-in.
`VADParams.stop_secs`0.20.2, raised to 0.6 to 0.8 only during capturePipecat's STT p99 constants assume 0.2. Changing it logs a warning.
`VADParams.min_volume`0.6Measure caller loudness firstSee the loudness math in the testing guide. PSTN and speakerphone callers often sit near the gate.
`SmartTurnParams.stop_secs`3.02.0 to 3.0This is the dead-air cost of a false "incomplete." Lower it for yes/no flows.
`user_turn_stop_timeout`5.05.0Backstop only. Watch for `on_user_turn_stop_timeout` firing; it means something upstream is broken.
`user_idle_timeout`0 (off)6 to 8 s, longer during captureCallers look up account numbers. See the idle section.
Start strategiesVAD + TranscriptionVAD + Transcription, or MinWords with escapeVAD start gives about 200 ms barge-in.
Audio input rate16,000Test 16,000 against 8,000At 8,000, Silero runs on native 8 kHz windows and Smart Turn resamples to 16 kHz internally. Neither has published phone results, so A/B test it.

Here is a runnable-shaped configuration for a Twilio media-streams agent using these settings:

# Illustrative configuration for pipecat-ai 1.12.0 on Twilio media streams.
from pipecat.audio.turn.smart_turn.base_smart_turn import SmartTurnParams
from pipecat.audio.turn.smart_turn.local_smart_turn_v3 import LocalSmartTurnAnalyzerV3
from pipecat.audio.vad.silero import SileroVADAnalyzer
from pipecat.audio.vad.vad_analyzer import VADParams
from pipecat.pipeline.worker import PipelineParams, PipelineWorker
from pipecat.processors.aggregators.llm_context import LLMContext
from pipecat.processors.aggregators.llm_response_universal import (
    LLMContextAggregatorPair,
    LLMUserAggregatorParams,
)
from pipecat.turns.user_mute import FunctionCallUserMuteStrategy
from pipecat.turns.user_start import TranscriptionUserTurnStartStrategy
from pipecat.turns.user_stop import TurnAnalyzerUserTurnStopStrategy
from pipecat.turns.user_turn_strategies import UserTurnStrategies

PHONE_VAD = VADParams(confidence=0.7, start_secs=0.2, stop_secs=0.2, min_volume=0.6)

context = LLMContext(messages)
user_aggregator, assistant_aggregator = LLMContextAggregatorPair(
    context,
    user_params=LLMUserAggregatorParams(
        vad_analyzer=SileroVADAnalyzer(params=PHONE_VAD),
        user_turn_strategies=UserTurnStrategies(
            start=[
                MinWordsWithEscapeStartStrategy(min_words=3),  # from the section above
                TranscriptionUserTurnStartStrategy(),
            ],
            stop=[
                TurnAnalyzerUserTurnStopStrategy(
                    turn_analyzer=LocalSmartTurnAnalyzerV3(
                        params=SmartTurnParams(stop_secs=2.5)
                    )
                )
            ],
        ),
        user_mute_strategies=[FunctionCallUserMuteStrategy()],
        user_idle_timeout=7.0,
    ),
)

worker = PipelineWorker(
    pipeline,  # transport.input(), stt, user_aggregator, llm, tts, transport.output(), assistant_aggregator
    params=PipelineParams(audio_in_sample_rate=16000, enable_metrics=True),
)

If your callers mostly interrupt with full sentences and you can tolerate occasional backchannel cutoffs, drop the custom class and keep the default VAD start. That is the lower-latency choice.

Noisy callers and speakerphones

Background noise causes two different failures, and they need different fixes.

  • False barge-ins (a TV, a coworker, road noise opening a turn while the bot speaks). VAD thresholds help only when the noise is quieter than the caller. Use noise suppression ahead of the VAD, which Pipecat supports through its audio filters, or a word-based start strategy. Our VAD misfire guide covers the taxonomy.
  • Late or missing end-of-turn (noise keeps VAD in "speaking," so no stop fires and Smart Turn never runs). The turn then ends only on the 5-second watchdog. Watch `on_user_turn_stop_timeout`. If it fires on more than a small fraction of turns, noise is holding VAD open.

Number and address readouts: a capture mode

Since Smart Turn cannot know the caller is reading digits, tell the pipeline. Pipecat has no public frame for swapping turn strategies mid-call. It does have `VADParamsUpdateFrame`, which the aggregator's VAD controller applies immediately. Pipecat's own IVR navigator uses the same frame to set `stop_secs=2.0` while it listens to phone menus.

Raising `stop_secs` to 0.7 during capture means the 300 to 500 ms pauses between digit groups no longer register as VAD stops, so Smart Turn is not asked to judge them. Barge-in is unaffected, because start timing is unchanged.

# Illustrative capture mode for pipecat-ai 1.12.0.
from pipecat.frames.frames import UserIdleTimeoutUpdateFrame, VADParamsUpdateFrame

CAPTURE_VAD = PHONE_VAD.model_copy(update={"stop_secs": 0.7})
capture_active = False

async def enter_capture_mode():
    global capture_active
    capture_active = True
    await worker.queue_frames([
        VADParamsUpdateFrame(params=CAPTURE_VAD),
        UserIdleTimeoutUpdateFrame(timeout=12.0),  # time to find a card
    ])

async def exit_capture_mode():
    global capture_active
    capture_active = False
    await worker.queue_frames([
        VADParamsUpdateFrame(params=PHONE_VAD),
        UserIdleTimeoutUpdateFrame(timeout=7.0),
    ])

# Enter from a function the LLM calls before asking for the number, for example
# async def request_callback_number(params): await enter_capture_mode(); await params.result_callback({"ok": True})

@user_aggregator.event_handler("on_user_turn_stopped")
async def on_user_turn_stopped(aggregator, strategy, message):
    digits = sum(ch.isdigit() for ch in (message.content or ""))
    if capture_active and digits >= 10:
        await exit_capture_mode()

Two side effects to plan for. The first change of `stop_secs` logs a warning that Pipecat's STT p99 constant assumed 0.2 s. And Deepgram's Finalize message is sent on VAD stop, so transcripts during capture finalize later. Both are acceptable for a few turns per call. If you would rather decide completeness semantically, `FilterIncompleteUserTurnStrategies` asks the LLM to prefix every reply with a completion marker. A partial phone number gets the "incomplete short" marker and a 5-second wait (`incomplete_short_timeout`). That costs an LLM call per detector trigger and is worth testing against capture mode.

User idle timeout on phone calls

`user_idle_timeout` replaced `UserIdleProcessor`. Read `UserIdleController` before you rely on it:

  • The timer starts on `BotStoppedSpeakingFrame`. If the bot has not spoken yet, for example on an outbound call where the bot waits for "hello," idle never arms.
  • It is suppressed while a user turn is open and while any function call is pending. A slow CRM lookup does not trigger "Are you still there?"
  • It cancels on `UserStartedSpeakingFrame` or `BotStartedSpeakingFrame`. With MinWords as the only start strategy, a mumble that never transcribes does not cancel it.
  • `UserIdleTimeoutUpdateFrame` applies immediately and can re-arm a running timer, which is how the capture-mode code extends it.

The Pipecat example `turn-management-detect-user-idle.py` shows the escalation pattern: a gentle prompt, a direct prompt, then a goodbye `TTSSpeakFrame` and an `EndWorkerFrame`.

Decision matrix: min-words, Smart Turn, or both

SituationStart strategyStop strategyExtra
General support, open questionsVAD + Transcription (default)Smart TurnDefaults are a good baseline
Bot reads long policies, lists, or confirmationsMinWords with escapeSmart Turn (1.7.0+)Keyword escape is mandatory
Number, address, or card captureUnchangedSmart Turn + capture mode, or FilterIncompleteLonger idle timeout during capture
Required legal disclosure at call startAnySmart Turn`MuteUntilFirstBotCompleteUserMuteStrategy`
Tool calls the caller should not cut offAnySmart Turn`FunctionCallUserMuteStrategy`
Heavy background speechMinWords with escape, or Krisp IPSmart TurnNoise filter before VAD
Spanish or mixed-language callersDefaultSmart TurnReport premature rate per language
Realtime speech-to-speech LLMDefaultSmart Turn with `wait_for_transcript=False``LLMContextAggregatorPair` sets this when `realtime_service_mode` is on
Need strict silence-based timingDefault`SpeechTimeoutUserTurnStopStrategy`Expect about 450 ms more per turn than Smart Turn

The rule underneath: start strategies trade barge-in speed against backchannel robustness, and stop strategies trade response speed against premature cutoffs. MinWords and Smart Turn solve different halves of the problem, so "both" is often the right answer. The endpointing comparison covers provider-side alternatives such as STT services with built-in end-of-turn detection.

Gotchas collected from the source and changelog

GotchaWhat happensFix
Old tutorials use `allow_interruptions`Silently ignored by the pydantic `PipelineParams` modelUse `user_turn_strategies`
VAD passed to `TransportParams`Silently ignored since 1.0; no VAD at allPass `vad_analyzer` to `LLMUserAggregatorParams`
MinWords counts per transcript framePaused "no... wait... stop" never interruptsKeyword escape
MinWords discards short utterancesCaller's "no" never reaches the LLMKeyword escape, or log resets
MinWords + Smart Turn before 1.7.0Bot answers each transcript fragmentUpgrade to 1.7.0 or later
Smart Turn sees only the last 8 sLong turns judged on their tailExpected; test long turns
Changing `stop_secs`Warning; STT p99 assumption breaksRe-run stt-benchmark, or change only during capture
`audio_idle_timeout=1.0`VAD force-stops if audio frames stop arriving mid-speechCheck carriers that stop sending media on mute or hold
Passing `ExternalUserTurnStrategies()` by handOverrides a service's `should_interrupt=False`Use `ExternalUserTurnStrategies(enable_interruptions=False)`
Idle never fires on outbound callsTimer arms only after the bot speaksHave the bot speak first, or queue `UserIdleTimeoutUpdateFrame`

How to test Pipecat turn settings before callers do

Use this protocol each time you change a turn setting, the STT provider, or the Pipecat version.

1. Instrument the pipeline. Add an observer that records turn frames and Smart Turn verdicts, and enable metrics so `TurnMetricsData` is emitted:

# Illustrative observer for pipecat-ai 1.12.0.
from pipecat.frames.frames import (
    BotStartedSpeakingFrame, BotStoppedSpeakingFrame, InterruptionFrame, MetricsFrame,
    UserStartedSpeakingFrame, UserStoppedSpeakingFrame,
    VADUserStartedSpeakingFrame, VADUserStoppedSpeakingFrame,
)
from pipecat.metrics.metrics import TurnMetricsData
from pipecat.observers.base_observer import BaseObserver, FramePushed
from pipecat.processors.frame_processor import FrameDirection

WATCH = (VADUserStartedSpeakingFrame, VADUserStoppedSpeakingFrame,
         UserStartedSpeakingFrame, UserStoppedSpeakingFrame,
         BotStartedSpeakingFrame, BotStoppedSpeakingFrame, InterruptionFrame)

class TurnEventLog(BaseObserver):
    def __init__(self):
        super().__init__()
        self.events = []  # (seconds, name, extra)

    async def on_push_frame(self, data: FramePushed):
        if not data.first_push or data.direction != FrameDirection.DOWNSTREAM:
            return
        t = data.timestamp / 1e9
        if isinstance(data.frame, WATCH):
            self.events.append((t, type(data.frame).__name__, None))
        elif isinstance(data.frame, MetricsFrame):
            for m in data.frame.data:
                if isinstance(m, TurnMetricsData):
                    self.events.append((t, "SmartTurn", (m.is_complete, m.probability,
                                                         m.e2e_processing_time_ms)))

turn_log = TurnEventLog()
worker = PipelineWorker(pipeline, params=PipelineParams(enable_metrics=True),
                        observers=[turn_log])

2. Compute the three core metrics from the log.

# Simplified metric extraction from TurnEventLog.events.
def turn_metrics(events, resume_window=2.0, barge_window=1.0):
    ev = sorted(events, key=lambda e: e[0])
    premature = releases = 0
    barge_ok = barge_total = 0
    bot_speaking = False
    for i, (t, name, _) in enumerate(ev):
        if name == "UserStoppedSpeakingFrame":
            releases += 1
            # Caller resumes before the bot has answered: a premature release.
            for t2, n2, _ in ev[i + 1:]:
                if t2 - t > resume_window or n2 == "BotStartedSpeakingFrame":
                    break
                if n2 == "VADUserStartedSpeakingFrame":
                    premature += 1
                    break
        elif name == "BotStartedSpeakingFrame":
            bot_speaking = True
        elif name == "BotStoppedSpeakingFrame":
            bot_speaking = False
        elif name == "VADUserStartedSpeakingFrame" and bot_speaking:
            barge_total += 1  # label each one barge-in or backchannel from the script
            if any(n2 == "InterruptionFrame" and 0 <= t2 - t <= barge_window
                   for t2, n2, _ in ev[i + 1:i + 40]):
                barge_ok += 1
    return {"premature_turn_rate": premature / max(releases, 1),
            "speech_while_bot_interrupted": barge_ok / max(barge_total, 1)}

For response latency, use the built-in `UserBotLatencyObserver` and its `on_latency_measured(observer, latency)` event. It measures from VAD stop, corrected for `stop_secs`, to `BotStartedSpeakingFrame`. Report p50 and p95, never the mean.

3. Build the test set. Eight scenario families, each run under four audio conditions, all rendered through the real 8 kHz μ-law path rather than WebRTC. Our SIP vs WebRTC comparison explains why the paths differ.

FamilyExampleMetricStarting pass threshold
S1 Short complete answers"Yes," "Tuesday at three"p95 response latency1.6 s or less before carrier
S2 Mid-turn hesitation"I need to, um, move my appointment"Premature turn rate3% or less clean, 5% or less noisy
S3 Phone number in groups"Five five five, eight six seven, five three oh nine"Captured in one turn95% or more
S4 Address with pausesStreet, unit, city, ZIPCaptured in one turn90% or more
S5 Backchannels during bot speech"Mm-hm," "okay," "right"False barge-in rate5% or less
S6 One-word barge-ins"No," "stop," "wait"Missed barge-in rate5% or less
S7 Sentence barge-ins"No, that's the wrong address"Missed barge-in rate; barge-in latency2% or less; p95 under 1.0 s
S8 Silence after a questionCaller says nothingIdle prompt fires on timeWithin 0.5 s of `user_idle_timeout`, every run

Conditions: clean, office babble at 10 dB SNR, competing TV speech at 10 dB SNR, and low-level speakerphone (about 12 dB quieter). Add a Spanish copy of S2 and S3 if you serve Spanish callers. These thresholds are our recommended starting targets, not industry standards; tighten them once you have a baseline. Our guide to synthetic callers covers how to render these turns, and background-noise testing covers SNR mixing.

4. Size the runs. To estimate a premature rate near 5% within plus or minus 3 points at 95% confidence, you need `n = 1.96^2 × 0.05 × 0.95 / 0.03^2 ≈ 203` turns per family. Across four conditions, that is about 51 turns per cell. To compare two configurations, run both on the same audio files so differences come from the setting, not the sample.

5. Read Smart Turn probabilities, not just verdicts. From the `SmartTurn` events, plot the probability of decisions made at pauses that turned out to be mid-turn. A cluster between 0.5 and 0.65 means premature cutoffs that a slightly different model or a capture mode would fix. A false "incomplete" cluster just under 0.5 on S1 explains 3-second dead air.

6. Gate releases on regression. Re-run S2, S3, S6 and S7 on every Pipecat upgrade. The changelog shows turn-handling fixes in almost every minor release from 1.5 to 1.12, and each one changes timing. Our interruption detection evaluation guide covers scoring when you also want human labels.

Where independent evaluation fits

Everything above can be run in-house, and it should be, at least once. The hard part is keeping it running: rendering 800 phone-quality calls per change, labeling which overlaps were real interruptions, and comparing configurations on identical audio. Evalgent does that as an independent third party. It drives scripted callers through your real phone number, scores premature cutoffs, missed barge-ins and latency per scenario, and keeps the same suite for regression runs after every Pipecat upgrade or provider swap. It fits at four points: a pre-launch audit, a bake-off between turn strategies or STT providers, regression runs after upgrades, and ongoing scoring of production calls.

Frequently asked questions

Does Pipecat still support allow_interruptions?

No. `allow_interruptions` was deprecated in 0.0.99 and removed in 1.0.0. Interruptions are on by default. To control them, pass `user_turn_strategies` to `LLMUserAggregatorParams`. Each start strategy has an `enable_interruptions` flag, and mute strategies such as `FunctionCallUserMuteStrategy` block interruptions at specific moments.

What replaced MinWordsInterruptionStrategy in Pipecat?

`MinWordsUserTurnStartStrategy` in `pipecat.turns.user_start`, placed in `UserTurnStrategies(start=[...])`. It applies the word minimum only while the bot is speaking and counts words per transcription frame. Below-threshold words are discarded, so add a keyword escape so that "no," "stop," and "wait" still interrupt.

How does Smart Turn v3 work with Silero VAD?

Silero detects pauses using 32 ms windows. Each time VAD reports a stop, which takes 200 ms of silence by default, Smart Turn scores up to the last 8 seconds of turn audio at 16 kHz. Above 0.5 probability, the turn releases once a finalized transcript arrives. Otherwise it waits for more speech or a 3-second silence fallback.

What VAD settings should I use for Twilio phone calls in Pipecat?

Start with the 1.12.0 defaults: confidence 0.7, start 0.2 s, stop 0.2 s, min_volume 0.6. Measure caller loudness before adjusting min_volume, and consider start_secs 0.3 on noisy lines. Raise stop_secs to about 0.7 only while callers read numbers. Validate every change over the real 8 kHz path.

Why does my Pipecat agent cut callers off while they read a phone number?

Each pause between digit groups is a VAD stop, and each stop is a separate Smart Turn decision. Small per-decision error rates compound across five or more pauses. The model has no number-entry mode. Use a capture mode that raises VAD stop_secs during the readout, or LLM-based completion filtering.

How much latency does Smart Turn add in Pipecat?

On the happy path, almost none. Inference takes about 65 ms on a Pipecat Cloud 1x instance and runs while the STT is still finalizing the transcript. The real cost is a false "incomplete" verdict, which waits for Smart Turn's stop_secs fallback, 3 seconds by default, before the bot replies.

Why doesn't user_idle_timeout fire in my Pipecat agent?

The idle timer starts only after `BotStoppedSpeakingFrame`. It is suppressed during open user turns and pending function calls, and it is cancelled when anyone starts speaking. If your bot never speaks first, or a tool call is still running, idle never arms. Queue `UserIdleTimeoutUpdateFrame` to arm or change it at runtime.

Should I use MinWords, Smart Turn, or both?

They solve different problems. MinWords filters backchannels at the start of a caller turn, while Smart Turn decides when the turn ends. Use both when the bot reads long content and callers say "mm-hm," on Pipecat 1.7.0 or later, with a keyword escape. Keep the default VAD start when fast barge-in matters most.

The bottom line

In Pipecat 1.x, interruption and turn-taking behavior comes from three settings in the user aggregator: the VAD timers, the start strategy and the stop strategy. On phone audio, each one needs a deliberate choice and a measurement. Start from the defaults, add a keyword escape and a number-capture mode, and test the eight scenario families on real 8 kHz calls before and after every change.

Related Articles