Evalgent
Back to Blog
Testing Strategies

LiveKit voice agent testing guide: unit tests, simulations, SIP and load (2026)

Deepesh Jayal
Updated
25 min read
LiveKit voice agent testing guide: unit tests, simulations, SIP and load (2026)
On this page

LiveKit now ships a full testing toolkit inside the framework: pytest helpers, eight built-in LLM judges, goal-driven simulations in text and audio modes, and a CLI load tester. Most teams wire up the first two and assume they are covered.

They are not. The layers that run on every commit exercise your LLM and tools. They do not run your VAD, your STT, your turn detector or your noise cancellation, and those cause most production complaints. Worse, several of those components behave differently in `dev` mode, on LiveKit Cloud and on a self-hosted agent server. So a green test run on a laptop can describe a different agent than the one answering calls.

This guide maps each LiveKit testing tool to the part of the pipeline it actually exercises. It breaks a LiveKit turn down hop by hop with a millisecond budget and explains what each metric field measures. Then it lays out a five-layer test plan with code and thresholds, plus the gotchas buried in docs footnotes and GitHub issues.

32 ms
VAD audio window the agent process must clear in real time (LiveKit)
0.3 s / 2.5 s
Default min/max endpointing delay with the audio turn detector (LiveKit docs)
84% to 52%
Turn-taking model hold/shift accuracy, clean vs 10 dB music noise (Russell and Harte, Interspeech 2025)
0-200 ms
Modal gap between human turns across 10 languages (Stivers et al., PNAS 2009)

What LiveKit's built-in testing covers, and what it skips

LiveKit's Test and Evaluate docs describe a lifecycle: Agent Console for manual debugging, unit tests per commit, simulations for whole conversations, and observability after release. The table below shows which pipeline stages each tool actually runs.

ToolWhat runsWhat does not runCost driverSensible cadence
Unit tests (`session.run()` + `judge()`)LLM, tools, handoffs, chat contextRoom connection, VAD, STT, turn detector, TTS, noise cancellationLLM tokensEvery commit
`JudgeGroup` with built-in judgesLLM grading of full `session.history`Anything audioJudge tokensEvery commit
Agent Simulations, text modeYour real entrypoint, LLM, tools, simulated userSTT, TTS, VAD and audio I/O are disabled automaticallySimulator, agent and judge tokensEvery commit
Agent Simulations, audio modeFull STT-LLM-TTS pipeline over a real audio track, with optional noise, low-quality mic and packet lossPSTN codecs, your SIP trunkReal-time runs, STT/TTS billing, audio turnsNightly or pre-release
`lk perf agent-load-test`Dispatch, room join, agent greeting, echo audioRealistic caller speechAgent minutes, inference quotasPre-release, infra changes
Real SIP test callsTrunk, dispatch rule, codec path, DTMF, hangupsScaleTelephony minutesPre-release, trunk changes

Two lines in LiveKit's own docs define the gap. The unit-testing page says tests make no LiveKit room connection. The simulations page says a text simulation turns off STT, TTS, VAD and audio input and output automatically.

So if your CI runs unit tests plus text simulations, nothing that runs on every merge touches turn-taking. That is a sensible cost trade-off. It just means audio-mode simulations and real calls carry the whole weight of turn detection, interruption handling and transcription accuracy.

Anatomy of a LiveKit turn, in milliseconds

To test turn-taking you need to know what happens between "the caller stops talking" and "the caller hears a reply." Here is the path through a cascaded LiveKit agent, with the parameter that governs each hop.

1. Audio arrives. A WebRTC caller sends Opus. A phone caller arrives as a SIP participant, usually on narrowband G.711 at 8 kHz. LiveKit's HD voice page notes that wideband G.722 needs a provider that supports it. Among third-party trunks that is currently Telnyx, and LiveKit Phone Numbers supports it out of the box.

2. Noise cancellation (optional). If you set `noise_cancellation` in `room_io.AudioInputOptions`, it runs before VAD, STT and the turn detector. Every downstream signal sees the processed audio.

3. VAD. The bundled inference VAD resamples input to 16 kHz and scores 512-sample windows, which is 32 ms of audio per window. In the current source, speech starts after `min_speech_duration` (0.05 s) above `activation_threshold` (0.5). It ends after `min_silence_duration` of silence: 0.25 s by default for the inference VAD and 0.55 s for the older Silero plugin. When END_OF_SPEECH fires, the framework back-dates the user's "stopped speaking" anchor by the silence duration plus inference time.

4. STT. Streaming STT emits interim and final transcripts. LiveKit's short-utterances guide is explicit: closing a turn requires both a final transcript and an end-of-speech signal. Interim transcripts never close a turn.

5. Turn detector. The default `inference.TurnDetector()` is an audio model that scores whether the turn is finished. If it doesn't return within about a second, the agent commits the turn anyway.

6. Endpointing wait. If the detector says "complete," the session waits `min_delay`. If it says "incomplete," it waits up to `max_delay`. Both are measured from the back-dated end of speech, so VAD silence counts toward the wait. With the audio turn detector the defaults are 0.3 s and 2.5 s.

7. Your hook, then the LLM. `on_user_turn_completed` runs (RAG lookups often live here), then the LLM streams tokens. Preemptive generation, on by default, can start the LLM when the final transcript lands, before the turn is confirmed.

8. TTS and playout. TTS starts on the first text, returns its first audio chunk, and the frames travel back to the caller.

Waterfall chart of one LiveKit voice agent turn showing VAD, STT final transcript, turn detector, endpointing wait, LLM time to first token, TTS first byte and transport, with agent-side and caller-heard totals

A worked latency budget

The numbers below are illustrative assumptions for a well-tuned agent co-located with its models. They are not measurements. Replace them with your own p50 values from `ChatMessage.metrics`.

HopGoverned byWebRTC (assumed ms)SIP/PSTN (assumed ms)LiveKit field
Uplink to agentNetwork, jitter buffer, SIP leg60150Not measured by the agent
VAD end of speech`min_silence_duration`250 (absorbed below)250 (absorbed below)Back-dated into the anchor
Final transcriptSTT provider endpointing200250`transcription_delay`
Endpointing wait (detector says complete)`min_delay` = 300300300`end_of_turn_delay`
Your hook`on_user_turn_completed`2020`on_user_turn_completed_delay`
LLM first tokenModel, prompt size, region350350`llm_node_ttft`
TTS first audioVoice model, region150150`tts_node_ttfb`
Downlink playoutNetwork, SIP leg60150Not measured by the agent

LiveKit's own approximation is `end_of_utterance_delay + llm.ttft + tts.ttfb`. Plugging in the assumptions:

  • Agent-side e2e: 300 + 20 + 350 + 150 = 820 ms. This is roughly what `e2e_latency` reports.
  • Caller-heard, WebRTC: 60 + 820 + 60 = 940 ms.
  • Caller-heard, SIP: 150 + 820 + 150 = 1,120 ms.
  • Same turn when the detector says "incomplete": the 300 ms wait becomes 2,500 ms. Agent-side e2e jumps to 3,020 ms.

Two lessons fall out of this arithmetic.

First, the end-of-turn delay is bimodal. Finished turns cluster near `min_delay` when the detector is confident and near `max_delay` when it is not. A healthy p50 can hide an ugly p90 made almost entirely of "detector unsure" turns. Track the max-delay rate: the share of committed user turns whose `end_of_turn_delay` is at least 80% of `max_delay`. That single number moves when a prompt change, a new STT model or a noise profile confuses the detector.

Second, `e2e_latency` is agent-side. LiveKit's simulation docs say it directly: the agent measures when it started producing audio, the caller measures when they heard it, and the difference is what the user perceives. On phone calls the uplink and downlink legs are invisible to the agent's own metrics. Audio simulations report caller-heard latency separately for this reason.

For context, human conversation sets a hard target. Stivers et al. (PNAS 2009) measured turn transitions in 10 languages. Every language showed a unimodal distribution with its mode between 0 and 200 ms. Means ranged from 7 ms in Japanese to 469 ms in Danish. A cascaded pipeline with a 300 ms minimum wait cannot match that, which is why preemptive generation and confident end-of-turn prediction matter more than shaving 20 ms off TTS.

What each LiveKit metric actually measures

LiveKit exposes four metric surfaces: per-plugin `metrics_collected` events, per-turn `ChatMessage.metrics`, live `session.usage`, and the end-of-session `SessionReport`. The session-level `metrics_collected` event is deprecated. The table maps fields to definitions from the data hooks docs.

FieldWhereWhat it measuresTesting trap
`transcription_delay`User messageEnd of speech to final transcriptClamped at 0, so a transcript that arrives "before" the anchor reads as 0
`end_of_turn_delay`User messageEnd of speech to the decision to end the turnIncludes VAD silence and transcription wait; bimodal
`end_of_utterance_delay``EOUMetrics`Same span, emitted per plugin eventNot emitted with server-side realtime turn detection
`on_user_turn_completed_delay`User messageTime inside your hookYour RAG latency hides here
`llm_node_ttft`Assistant messageLLM first tokenEmpty for realtime models
`tts_node_ttfb`Assistant messageFirst audio chunk after the first text tokenEmpty for realtime models
`playback_latency`Assistant messageFirst frame forwarded to playback startedNear zero unless an avatar worker reports playback
`e2e_latency`Assistant messageUser stopped speaking to agent began respondingAgent-side only, excludes network legs
`detection_delay``InterruptionMetrics`Overlap onset to the adaptive model's final predictionOnly exists when adaptive interruption is active
`inference_duration_total` / `inference_count``VADMetrics`VAD compute per windowA rising average means CPU starvation

The metric-anchor bug you should guard against

GitHub issue #6093, filed against livekit-agents 1.6.0, documents turns with `transcription_delay` and `end_of_turn_delay` above 200 seconds. Other turns in the same sessions had no values at all. The cause was a stale "stopped speaking" anchor: when the turn detector split one long utterance into consecutive user turns, the second turn inherited an anchor from minutes earlier. The issue links two older reports, #2361 and #4388.

The current source now refuses to compute these metrics when the anchor predates the turn's start, and returns `None` instead. But the trace span attribute is written as `metrics.end_of_turn_delay or 0`. Turns without a valid metric therefore show up as 0 in traces, which drags averages down.

Three rules follow for any latency test or dashboard built on LiveKit metrics:

  • Track metric coverage. Coverage is committed user turns with a non-null `end_of_turn_delay` divided by all committed user turns. Treat a drop in coverage as a regression in its own right.
  • Drop exact zeros and values above `max_delay` plus 5 seconds from latency percentiles. Count them separately rather than averaging them in.
  • Correlate by `speech_id`. Greetings and `say()` calls have no `speech_id`, so exclude them from turn latency.

The environment trap: same code, different agent

The single most useful thing to know about LiveKit testing is that three components change with the environment, even if your code doesn't.

Component`dev` with LiveKit Cloud credentialsDeployed to LiveKit CloudSelf-hosted with `start`
Turn detector default`v1` full model, free monthly allowance, then falls back to `v1-mini``v1` full model`v1-mini` on local CPU
Interruption mode defaultAdaptive (40,000 free requests per month locally)AdaptiveVAD-based
Prewarmed idle processes0Managed`ceil(cpu_count)` in Python
Load functionn/aFixed, not configurableAverage CPU over 5 s, threshold 0.7
SIGTERM handlingNot reliable (LiveKit load-test guide)Graceful drainGraceful drain, `drain_timeout` default 1 hour

These come from the turn detector, adaptive interruption and server options pages. The turn detector page adds two details that matter for testing.

The fallback is sticky. If the full `v1` model times out or can't be reached, the session logs one warning and emits a default probability of 1.0 for the in-flight prediction. That means "turn complete," so the agent replies. It then runs `v1-mini` for the rest of the session. One network blip early in a call changes turn-taking for every remaining turn.

`v1-mini` shares your CPU. LiveKit recommends compute-optimized instances such as AWS c6i or c7i over burstable t3 or t4g, which can time out on CPU credits even when utilization looks low.

The practical rule is to pin what you test. Set `inference.TurnDetector(version=...)` explicitly and set `interruption.mode` explicitly. Then add a test that fails if those values drift. If you self-host on `v1-mini`, run your audio tests against `v1-mini`, not the `v1` model your laptop used in `dev`.

The audio simulation CLI already protects you here. The docs note that it uses the same turn detection and adaptive interruption defaults as a deployed agent, instead of falling back to local defaults.

A layered LiveKit test plan

The plan below is a framework you can copy. Each layer has a trigger, a tool, and a gate. The thresholds are suggested starting points, not industry standards. Tune them to your call types.

Five-layer LiveKit test plan stacked from unit tests to production scoring, showing tool, trigger, pipeline coverage and pass gate for each layer
LayerToolTriggerCoversSuggested gate
L1 Unitpytest + `session.run()`, `mock_tools`, `JudgeGroup`Every commitTool arguments, error paths, handoffs, refusals100% of deterministic asserts; judge checks pass 3 of 3 runs
L2 Conversation`lk agent simulate text`, then `audio` with degradation flagsText per commit, audio nightlyWhole-call outcomes; in audio mode, turn-taking, interruptions, entity accuracyNo scenario regresses from pass to fail; audio p95 caller-heard latency within 15% of baseline
L3 TelephonyReal calls through your trunk or LiveKit Phone NumbersPre-release, trunk or codec changeDispatch rule, codec, noise cancellation on SIP audio, DTMF, hangupsEvery SIP check in the checklist passes; entity accuracy on PSTN within 3 points of WebRTC
L4 Load`lk perf agent-load-test`, rampedPre-release, infra changeDispatch, join delay, CPU headroom, VAD keeping paceJoin delay p95 stable across ramp steps; no persistent "slower than realtime" warnings
L5 Production`on_session_end`, session reports, judges, external scoringEvery callReal callers, drift, regressions after model updatesAlert on max-delay rate, metric coverage, task completion

Layer 1: unit tests that hold up

The LiveKit test helpers are good. Use them for what they test well: deterministic checks on tool calls and text behavior. The example below is illustrative and simplified, but every class, method and parameter name comes from the current unit-testing docs.

# tests/test_booking.py  (illustrative; pytest + pytest-asyncio)
import json
import pytest
from livekit.agents import AgentSession, inference, mock_tools
from livekit.agents.evals import (
    Judge, JudgeGroup, JudgmentResult, task_completion_judge, tool_use_judge,
)
from agent import Assistant  # your Agent subclass

AGENT_MODEL = "openai/gpt-4.1-mini"     # the model you actually ship
JUDGE_MODEL = "google/gemma-4-31b-it"   # any LLM; need not match the agent's

@pytest.mark.asyncio
async def test_booking_extracts_normalized_arguments() -> None:
    async with (
        inference.LLM(model=AGENT_MODEL) as agent_llm,
        inference.LLM(model=JUDGE_MODEL) as judge_llm,
        AgentSession(llm=agent_llm) as session,
    ):
        await session.start(Assistant())
        result = await session.run(
            user_input="Book me Tuesday the 13th at three pm, customer ID 4471"
        )
        call = result.expect.next_event().is_function_call(name="book_appointment")
        args = json.loads(call.event().item.arguments)
        assert args["customer_id"] == "4471"   # assert normalized values,
        assert args["time"] in ("15:00", "3:00 PM", "3 pm")  # not one exact string
        result.expect.next_event().is_function_call_output()
        await result.expect.next_event().is_message(role="assistant").judge(
            judge_llm, intent="Confirms an appointment on Tuesday the 13th at 3 pm."
        )
        result.expect.no_more_events()

@pytest.mark.asyncio
async def test_booking_failure_is_not_confirmed() -> None:
    async with (
        inference.LLM(model=AGENT_MODEL) as agent_llm,
        inference.LLM(model=JUDGE_MODEL) as judge_llm,
        AgentSession(llm=agent_llm) as session,
    ):
        await session.start(Assistant())
        with mock_tools(
            Assistant, {"book_appointment": lambda: RuntimeError("booking API 503")}
        ):
            result = await session.run(user_input="Book me Tuesday at 3 pm, ID 4471")
            result.expect.next_event().is_function_call(name="book_appointment")
            out = result.expect.next_event().is_function_call_output()
            assert out.event().item.is_error
            await result.expect.next_event(type="message").judge(
                judge_llm,
                intent="Says the booking did not go through and offers a next step. "
                       "Does not claim the appointment is booked.",
            )

class NoConfirmationAfterToolError(Judge):
    """Deterministic: no 'confirmed' after a failed tool call."""
    def __init__(self) -> None:
        super().__init__(name="no_confirmation_after_tool_error")

    async def evaluate(self, *, chat_ctx, reference=None, llm=None) -> JudgmentResult:
        tool_failed = False
        for item in chat_ctx.items:
            if item.type == "function_call_output" and item.is_error:
                tool_failed = True
            elif tool_failed and item.type == "message" and item.role == "assistant":
                if "confirmed" in (item.text_content or "").lower():
                    return JudgmentResult(verdict="fail",
                                          reasoning="Claimed success after a tool error")
        return JudgmentResult(verdict="pass", reasoning="No false confirmation")

@pytest.mark.asyncio
async def test_whole_conversation() -> None:
    async with (
        inference.LLM(model=AGENT_MODEL) as agent_llm,
        AgentSession(llm=agent_llm) as session,
    ):
        await session.start(Assistant())
        await session.run(user_input="Hi, I need to move my appointment")
        await session.run(user_input="Make it Thursday at 10, ID 4471")
        verdict = await JudgeGroup(
            llm=JUDGE_MODEL,
            judges=[task_completion_judge(), tool_use_judge(),
                    NoConfirmationAfterToolError()],
        ).evaluate(session.history)
        assert verdict.all_passed, {k: v.reasoning for k, v in verdict.judgments.items()}

Add one cheap test that protects you from the environment trap and from unit mistakes. Python takes seconds and Node.js takes milliseconds. A `min_delay` of 500 copied from a Node.js config means 500 seconds in Python.

# tests/test_config.py  (illustrative; agent_config.py holds plain dicts you also pass to TurnHandlingOptions)
from agent_config import TURN_DETECTOR_VERSION, TURN_HANDLING

def test_turn_handling_is_pinned_and_sane() -> None:
    assert TURN_DETECTOR_VERSION in ("v1", "v1-mini")       # never left to auto-select
    ep = TURN_HANDLING["endpointing"]
    assert 0.1 <= ep["min_delay"] <= ep["max_delay"] <= 5.0  # seconds, not ms
    assert TURN_HANDLING["interruption"]["mode"] in ("adaptive", "vad")
    assert TURN_HANDLING["interruption"].get("min_words", 0) == 0  # see gotchas

Three unit-test traps from the docs

  • `judge()` sees one message. The docs say the judge evaluates the message "without surrounding conversation context." Anything that depends on earlier turns belongs in a `JudgeGroup` run over `session.history`.
  • `get_job_context()` raises `RuntimeError` in tests. Code paths that touch the job context need a mock, or they need to be tested elsewhere.
  • Judges are probabilistic. Our post on LLM-as-judge limits covers why. Put the hard checks in deterministic asserts and custom `Judge` subclasses, and keep LLM judges for intent.

How many green runs prove anything?

A flaky judge-based test that passes 10 times in a row tells you less than it feels like. The rule of three gives a quick bound: with zero failures in n independent runs, the 95% upper confidence bound on the failure rate is about 3/n.

  • 10 clean runs: the true failure rate could still be 30%.
  • 60 clean runs: about 5%.
  • 150 clean runs: about 2%.

Comparing two configurations needs more. To detect a scenario pass rate moving from 90% to 95% at 5% significance and 80% power, the standard two-proportion formula gives n = (1.96 + 0.84)² × (0.09 + 0.0475) / 0.05² = 7.84 × 0.1375 / 0.0025 ≈ 431 runs per arm. That is why A/B decisions on turn-taking belong in batched simulations and production data, not in CI. Our guide to A/B testing voice agents covers the design.

Layer 2: simulations, text first and audio second

Agent Simulations are in beta. They need CLI v2.16.4 or later for Python (v2.18.3 for audio), LiveKit Agents 1.6.6 or later, and a LiveKit Cloud project. A simulated user follows a scenario's `instructions`, your real agent runs locally under a temporary name, and a judge grades the transcript against `agent_expectations`.

The best feature is `on_simulation_end`. It lets you fail a run when the final state is wrong, even if the conversation read well. The final result is the logical AND of the judge verdict and your check, so a polished transcript that booked the wrong room still fails. Pair that with session-scoped `mock_tools(..., session=session)` seeded from `ctx.simulation_context()` userdata, and your tool flows become reproducible. We cover more patterns in detecting silent tool failures.

Write turn-taking scenarios on purpose. Most generated scenarios are clean transactions. The audio layer only earns its cost if the simulated caller does the things that break endpointing.

# scenarios_turn_taking.yaml  (illustrative; fields from LiveKit's scenario format)
name: Turn-taking regressions
scenarios:
  - label: Caller answers a yes/no question with a single word
    instructions: |
      PERSONA: Busy caller, answers in one word whenever possible.
      OPENING LINE: "Hi, I need to reschedule."
      DO, IN ORDER:
      1. When asked to confirm the existing appointment, say only "Yes."
      2. When offered Thursday at 10, say only "No."
      3. Accept the next slot offered. Do not hang up until it is confirmed.
    agent_expectations: >
      Responds to each one-word answer without the caller repeating it.
      Ends with exactly one rescheduled appointment. Any turn where the caller
      has to say "yes" or "no" twice is a fail.
    tags:
      feature: turn_taking
      pattern: short_utterance

Then run the same file in both modes, and add degradation in audio mode:

lk agent simulate text  --scenarios scenarios_turn_taking.yaml
lk agent simulate audio --scenarios scenarios_turn_taking.yaml --background-noise
lk agent simulate audio --scenarios scenarios_turn_taking.yaml --low-quality-microphone --packet-loss
lk agent simulate export <run-id> > run.json   # archive as a CI artifact

Audio runs report what text cannot: end-of-turn mispredictions, time to yield after a barge-in, false interruptions, unanswered caller turns, and word and entity error rates in both directions. Entity scoring separates "never recognized" from "recognized and later lost," which matters for confirmation codes. See STT entity accuracy for why entity error matters more than WER.

Plan wall-clock time. Audio runs execute in real time, with a default per-run concurrency of 15 and a per-project cap of 30. As a worked example, 120 audio scenarios averaging 2.5 minutes is 300 scenario-minutes. At 15 concurrent, the floor is 300 / 15 = 20 minutes, and it doubles if a second pipeline is using the project cap at the same time.

What the research says about noisy turn-taking

LiveKit's default turn detector is an audio model. It reads intonation and rhythm as well as words. That is a strength in clean audio and a test obligation in noisy audio.

  • Ekstedt and Skantze (Interspeech 2022) introduced Voice Activity Projection. It is a self-supervised model that predicts upcoming voice activity, which lets it anticipate turn shifts and backchannels without labeled data. It is the research lineage behind acoustic end-of-turn prediction.
  • Russell and Harte (Interspeech 2025) tested predictive turn-taking models in noise. Hold/shift accuracy fell from 84% in clean speech to 52% in 10 dB music noise, close to a coin flip. They also found that training relied on accurate transcription, which limited ASR-derived transcripts to clean conditions.
  • Full-Duplex-Bench (Lin et al., 2025) defines metrics you can reuse for any agent. Takeover rate during user pauses measures how often the system grabs the floor when it shouldn't. It also scores backchannel timing against human behavior, response latency on smooth turn-taking, and behavior under user interruption.

In practice: run your turn-taking scenarios at more than one noise level, and report max-delay rate and premature-takeover rate per condition. A detector that is excellent in a quiet room can be near chance in a car with the radio on. Our guide to testing STT under background noise covers building SNR-controlled audio.

Timeline of adaptive interruption handling in LiveKit showing the one-second start cooldown, a backchannel ignored mid-utterance, a true barge-in yielding, and false-interruption resume after two seconds

Interruption handling has a blind second

Adaptive interruption handling separates real barge-ins from backchannels such as "uh-huh." It needs LiveKit Cloud or `dev` mode, a VAD, and an STT that supports aligned transcripts. Check `stt.capabilities.aligned_transcript`. Otherwise the session falls back to VAD-based interruption.

The detail that changes your tests is `backchannel_boundary`, which defaults to `(1.0, 1.0)`. For the first second after the agent starts speaking, adaptive detection is suppressed and VAD-based interruption is used, so real corrections aren't swallowed. That means a backchannel such as "okay, sure" that lasts past `min_duration` (0.5 s) in the first second still interrupts the agent, while the same words three seconds in are ignored.

Test backchannels at two offsets, inside the first second of agent speech and well after it, and at two lengths, under and over 0.5 s. Also test the false-interruption path. With defaults, speech of at least 0.5 s (`min_duration`) pauses the agent. If no transcript arrives within `false_interruption_timeout` (2.0 s), `agent_false_interruption` fires and the agent resumes. A cough should produce one resume, not a restarted sentence. Our interruption detection guide has scoring definitions.

Layer 3: real SIP calls and phone audio

WebRTC tests flatter your agent. A phone caller reaches it through a trunk, a dispatch rule, a SIP participant and usually an 8 kHz codec. The inference VAD resamples that 8 kHz audio up to 16 kHz, but upsampling cannot restore the missing band above 4 kHz. Your STT and turn detector both work with less signal than they had in the browser. For the transport trade-offs, see SIP vs WebRTC for voice agents.

LiveKit's telephony testing page gives a checklist. Turned into tests, it looks like this:

CheckHowPass condition
Number and dispatch rule exist`lk number list` or `lk sip inbound list`, `lk sip dispatch list`Rule matches the trunk; `agent_name` matches your code
Agent is registeredHealth check on port 8081 (production default)HTTP 200, not 503
Correct rule matched`lk room list`Room name has the rule's prefix
SIP participant attributes`lk room participants get --room ``kind` is SIP; `sip.callID`, `sip.trunkID`, `sip.ruleID` as expected
Outbound failure paths`CreateSIPParticipant` with `wait_until_answered=True``SipCallError` caught for `USER_REJECTED` and `USER_UNAVAILABLE`
Hangup mid-callCaller hangs upDisconnect reason `CLIENT_INITIATED`; session closes
DTMFKeypad input in an IVR-style stepDigits captured (`GetDtmfTask` is the prebuilt option)
Codec and regionProvider logs, region pinningWideband where available; trunk region near the agent

Two LiveKit facts limit what you can test with LiveKit Phone Numbers. They don't support outbound calling, so outbound tests need a third-party trunk. And LiveKit's April 2026 latency guide said they supported US numbers only at the time.

Noise cancellation can hurt STT

Noise cancellation is a test variable, not a free win. Three sources point the same way.

  • Iwamoto et al. (Interspeech 2022) decomposed single-channel speech-enhancement errors into residual noise and artifacts. Artifacts, not leftover noise, were the main cause of ASR degradation. Simply mixing a scaled copy of the original signal back into the enhanced audio improved recognition.
  • LiveKit's noise cancellation page shows WER on one demo clip transcribed by Deepgram Nova-3. The original scored 117.6% because of background speech. Krisp VIVA scored 11.8%, ai-coustics QUAIL_VF_S 7.1%, and QUAIL_VF_L 14.3%. It is a single clip, but the larger model did worse on it. Model size does not predict your result.
  • LiveKit's short-utterances guide warns that noise cancellation can classify a short "yes" as noise, especially on SIP. It also says to remove stacked models, for example Krisp enabled on the trunk (`krisp_enabled`) plus an enhanced model in the agent. The noise cancellation docs add that models are trained on raw audio, so frontend and agent-side cancellation shouldn't both be on.

If your STT is Whisper-based, test silence and hold music too. Koenecke et al. (FAccT 2024) found that about 1% of Whisper transcriptions contained entire hallucinated phrases. Hallucinations were more frequent for speakers with longer non-vocal stretches. A hallucinated final transcript during hold music can open a phantom turn.

The protocol: run the same SIP audio set with noise cancellation off, with background suppression, and with the telephony-tuned voice isolation model. Score entity accuracy and short-utterance response rate for each. Keep whichever wins on your audio, even if it is "off."

Layer 4: load testing with the lk CLI

The agent-aware load tester is `lk perf agent-load-test`. It is separate from `lk load-test`, which tests WebRTC transport and isn't agent-aware. From LiveKit's load-testing guide:

lk perf agent-load-test \
  --rooms 10 \
  --agent-name my-agent \
  --echo-speech-delay 10s \
  --duration 5m

Know what this measures. Each room gets an echo participant that replays your agent's own audio after the delay. So your agent must greet first, or nothing happens. The STT is then transcribing your TTS voice, which is clean synthetic speech. This test measures dispatch, join delay, CPU headroom and pipeline behavior under concurrency. It does not measure accuracy on real callers.

The CLI ramps on its own: it creates a room, waits for the agent to join, then creates the next. LiveKit's guidance adds:

  • Run agents in `start` mode, not `dev`.
  • Test the deployed agent, so the Cloud dashboard records join latency percentiles.
  • Run from cloud VMs, not office bandwidth.
  • Raise file-descriptor and socket limits on the load generator.
  • For custom scripts, ramp gradually, for example 10, 25, 50, 75, then 100 sessions a minute apart, and hold at 100 for at least five minutes.

Capacity math

Size the test from Little's law: concurrent calls = arrival rate × average handle time. The numbers below are a worked example, not a benchmark.

  • Peak traffic of 900 calls per hour is 15 calls per minute.
  • With a 4-minute handle time, that is 15 × 4 = 60 concurrent calls.
  • Test at 2× peak: 120 concurrent sessions.
  • On self-hosted servers, the default `load_fnc` averages CPU over 5 seconds and stops accepting jobs above 0.7. If one session costs a server 5% CPU, each server accepts 0.7 / 0.05 = 14 sessions. 120 / 14 ≈ 8.6, so plan 9 servers plus one spare.

Watch the VAD while you ramp. The warning "inference is slower than realtime" means the VAD in a job process has fallen behind live audio. The reported `delay` is how far behind, not how long inference took. LiveKit's explainer lists the symptoms: late end of speech, talking over callers, missed interruptions and clipped greetings. The usual cause is a blocked event loop, which can happen at low CPU. A brief burst in the first seconds of a session is harmless. A persistent warning under load is a failed test. See concurrency failures in voice agents for related patterns.

Layer 5: production scoring and correlation

Every call after launch is test data. LiveKit gives you the hooks: `on_session_end`, `ctx.make_session_report()`, and `JudgeGroup`. When a `JudgeGroup` runs inside a job context, it tags the session with `lk.judge.:`, which then surfaces in LiveKit Cloud. The report functions run in your process, so they also work self-hosted.

One query sends people to this page every month: how do I correlate LiveKit room events with evaluation results? Capture join keys at the start of the call, then emit them with your scores at the end.

# agent.py  (illustrative)
import logging
from livekit.agents import AgentServer, JobContext
from livekit.agents.evals import JudgeGroup, task_completion_judge, tool_use_judge

logger = logging.getLogger("calls")
server = AgentServer()
CALL_KEYS: dict[str, dict] = {}

async def on_session_end(ctx: JobContext) -> None:
    session = ctx.primary_session
    report = ctx.make_session_report().to_dict()
    verdict = await JudgeGroup(
        llm="openai/gpt-4o-mini",
        judges=[task_completion_judge(), tool_use_judge()],
    ).evaluate(session.history)  # in a job context this also tags lk.judge.<name>
    logger.info("call_scored", extra={
        **CALL_KEYS.pop(ctx.job.id, {}),
        "score": verdict.score,
        "verdicts": {k: v.verdict for k, v in verdict.judgments.items()},
        "report": report,  # history, events, options: ship to your warehouse
    })

@server.rtc_session(agent_name="my-agent", on_session_end=on_session_end)
async def entrypoint(ctx: JobContext) -> None:
    await ctx.connect()
    caller = await ctx.wait_for_participant()
    CALL_KEYS[ctx.job.id] = {
        "room": ctx.room.name,
        "job_id": ctx.job.id,
        "sip_call_id": caller.attributes.get("sip.callID"),  # matches carrier logs
    }
    # ... build AgentSession and start it

The join keys, from coarse to fine:

KeyJoins to
Room name and job IDLiveKit Cloud sessions, agent logs, Agent Insights
`sip.callID` / `sip.callIDFull`SIP provider CDRs and PCAPs
`sip.twilio.callSid`Twilio console, on Twilio trunks
`speech_id`Per-turn EOU, LLM and TTS metrics
`lk.judge.:` tagsLiveKit Cloud session filters

Two notes. `on_session_end` is bounded by `session_end_timeout` (default 5 minutes), so push heavy scoring to a queue. And a session that went wrong can become a regression scenario through "Turn into a test" in the dashboard. That closes the loop back to Layer 2.

Where independent evaluation fits: LiveKit's judges are useful self-checks, but they run inside the stack you're grading. Evalgent scores calls from outside the stack, against your own scenarios and outcomes. It works the same way on a LiveKit agent as on a Vapi or Retell agent, which is what you need for a vendor bake-off or for checking regressions after a model update. For what to capture per call, see what to log on every voice agent call.

The LiveKit turn-taking test matrix

This is the artifact to copy into your test plan. Rows are caller behaviors that break turn-taking. Columns are conditions. Run every cell at least five times in audio mode. The thresholds are our suggested starting points, not published standards.

Caller behaviorConditions to runPrimary metricSuggested pass gate
One-word answer ("yes", "42")Clean, background noise, SIP G.711, noise cancellation on and offShort-utterance response rate (agent replies without a re-prompt)At least 98%
Mid-sentence pause ("from Dublin... to New York")Clean, noise, slow speakerPremature takeover rateAt most 3%
Backchannel inside the first second of agent speechClean, noiseUnwanted interruption rateRecord it; it will be higher by design
Backchannel after the first secondClean, noiseUnwanted interruption rateAt most 5%
True barge-in correction ("no, Tuesday")Clean, noise, packet lossTime to yield; correction capturedYield p95 under 500 ms; 100% captured
Spelled entity (email, code)Low-quality mic, SIPEntity accuracyWithin 3 points of clean WebRTC
Long monologue`user_turn_limit` set and unsetAgent cuts in at the limit100% when set
Silence or hold musicWhisper-based STT, noise cancellation onPhantom turns per minute0

Across all rows, report max-delay rate, metric coverage and caller-heard p95 latency per condition. Seven behaviors across four or five conditions gives about 30 cells, or around 150 audio runs at five each. That is a nightly budget, not a per-commit one.

LiveKit gotchas that cost teams a week

Each item below comes from LiveKit docs, source code, or GitHub issues.

1. A final transcript is mandatory. Some STTs drop very short speech and never send a final transcript, so the turn never closes and the agent stays silent. The opt-in `transcription_timeout` on `AgentSession` emits `user_transcription_timeout`, so you can ask the caller to repeat.

2. `min_words` can swallow answers. If the caller replies while the agent is still talking and the reply is shorter than `min_words`, the turn is not committed at all.

3. STT endpointing stacks. With `turn_detection="stt"`, `min_delay` is added after the provider's end-of-speech signal. LiveKit suggests setting it to 0.

4. VAD floor. The audio turn detector raises `ValueError` if VAD `min_silence_duration` is below 0.25 s. Moving from the Silero plugin (0.55 s default) to the bundled inference VAD (0.25 s) changes timing even if you changed nothing else.

5. Realtime models ignore most interruption options. With server-side turn detection, only `enabled` and `discard_audio_if_uninterruptible` apply, and `enabled=False` raises `ValueError` at session start.

6. `llm_node_ttft` and `tts_node_ttfb` are empty for realtime models. `RealtimeModelMetrics.ttft` can be -1 when no audio tokens were produced. Filter before averaging.

7. The `user_turn_limit` cut-in can't be interrupted. The default `on_user_turn_exceeded` reply uses `allow_interruptions=False`.

8. `max_delay` is not a cure for short utterances. LiveKit warns that cutting endpointing delays too far can attach a final transcript to the wrong turn.

9. Burstable instances time out the local turn detector. Use compute-optimized instances for `v1-mini`.

10. The free tier cold-starts. LiveKit's latency guide notes that Build-plan projects cold-start agents after all sessions end. Don't benchmark join latency on the free plan.

How to test a LiveKit voice agent before launch

1. Pin the environment. Set the turn detector `version`, `interruption.mode` and endpointing values explicitly, and add a config test that fails on drift or unit mix-ups.

2. Write L1 unit tests. Cover tool arguments, tool errors via `mock_tools`, handoffs and refusals. Use deterministic asserts first and `JudgeGroup` for whole-conversation intent.

3. Commit a scenario set. Use outcome-shaped `agent_expectations`, `on_simulation_end` state checks, and turn-taking scenarios for short answers, pauses, backchannels and corrections.

4. Run text simulations on every commit and audio simulations nightly. Add `--background-noise`, `--low-quality-microphone` and `--packet-loss`, and export runs as CI artifacts.

5. Measure the turn properly. Track max-delay rate, metric coverage, caller-heard p95 and `e2e_latency` separately, filtering anchor-bug values.

6. Place real SIP calls. Run the telephony checklist, compare entity accuracy on PSTN audio with WebRTC, and A/B noise cancellation settings on phone audio.

7. Load test in `start` mode. Use `lk perf agent-load-test` against the deployed agent, ramp to 2× peak from Little's law, and fail the run on persistent VAD "slower than realtime" warnings.

8. Wire production scoring before launch. Emit join keys and judge verdicts from `on_session_end`, route bad sessions back into the scenario set, and add independent scoring for every call.

Where LiveKit sits against Pipecat and Vapi for testing

If you're still choosing a framework, testability is a fair criterion. LiveKit's advantage is that the test helpers, simulations and observability are built in and share one data model. The cost is that several defaults depend on running on LiveKit Cloud. Our comparisons of Pipecat vs LiveKit, how to choose between them, latency differences and LiveKit vs Vapi go deeper. Whatever you pick, run the same scenarios against each candidate so the comparison is fair. That vendor bake-off is one of the places Evalgent is used most, because the scoring sits outside every stack being compared.

Frequently asked questions

Does LiveKit have built-in testing for voice agents?

Yes. LiveKit Agents includes pytest and Vitest helpers built on `AgentSession.run()`, with assertions for messages, tool calls and handoffs. It also ships eight built-in LLM judges via `JudgeGroup`, beta Agent Simulations in text and audio modes, and the `lk perf agent-load-test` CLI. Unit tests and text simulations skip the audio pipeline entirely.

Do LiveKit unit tests exercise turn detection?

No. LiveKit's docs state that testing doesn't make a room connection, and text simulations automatically disable STT, TTS, VAD and audio I/O. Turn detection, interruption handling and transcription accuracy are only exercised by audio-mode simulations, real WebRTC or SIP calls, and production traffic.

Why does my LiveKit agent behave differently in dev and production?

Defaults depend on environment. In `dev` with Cloud credentials, the turn detector uses the full `v1` model and adaptive interruption. A self-hosted agent started with `start` defaults to the local `v1-mini` model and VAD-based interruptions. Pin `TurnDetector(version=...)` and `interruption.mode` explicitly, and test the configuration you actually deploy.

What does e2e_latency measure in LiveKit?

It is the time from when the agent detected the user stopped speaking to when the agent began responding. It is agent-side, so it excludes the network legs to and from the caller, which can be large on SIP. LiveKit approximates it as end-of-utterance delay plus LLM time to first token plus TTS time to first byte.

What is a good end-of-turn delay for a LiveKit agent?

With the audio turn detector, defaults are 0.3 s minimum and 2.5 s maximum. Confident turns land near the minimum and unsure turns near the maximum, so the distribution is bimodal. Track the share of turns near `max_delay` rather than the average. For reference, human gaps cluster between 0 and 200 ms.

How do I load test a LiveKit voice agent?

Use `lk perf agent-load-test` with `--rooms`, `--agent-name`, `--echo-speech-delay` and `--duration` against a deployed agent running in `start` mode. The agent must greet first because the echo participant only replays what it hears. Ramp gradually, run from cloud VMs, and size the target with Little's law at 2× peak.

Should I enable noise cancellation for SIP callers?

Test it rather than assume it. Use one model, not stacked trunk and agent cancellation, and prefer the telephony-tuned variant for SIP audio. Research shows enhancement artifacts can raise ASR errors, and LiveKit warns that cancellation can suppress short replies. Compare entity accuracy and short-answer response rate with it on and off.

How do I correlate LiveKit room events with evaluation results?

Capture join keys when the call starts: room name, job ID, and the SIP participant's `sip.callID`. Emit them with judge verdicts and the session report from `on_session_end`. Use `speech_id` to join per-turn metrics. `JudgeGroup` running in a job context also tags sessions with `lk.judge.:` in LiveKit Cloud.

The bottom line

LiveKit gives you strong tools for testing text and tool logic, but turn-taking, transcription and phone audio only get tested when you run audio simulations, real SIP calls and ramped load tests against the configuration you actually deploy. Pin your environment, measure the turn hop by hop, and score every production call, and your test results will finally describe the agent your callers hear.

Related Articles