Evalgent
Back to Blog
Voice AI Testing

Debugging a LiveKit Latency Issue End to End: Metrics, Waterfalls and Fixes

Deepesh Jayal
23 min read
Debugging a LiveKit Latency Issue End to End: Metrics, Waterfalls and Fixes
On this page

Search "livekit latency issue" and you find the same thread written fifty times. Someone's STT, LLM and TTS numbers add up to under a second, yet callers wait three. One community thread is titled exactly that: "ASR+LLM+TTS total <1s, but speech end to reply takes 3-4s." A GitHub issue reports 3,000 to 4,000 ms through Twilio and 6,000 to 7,000 ms through Telnyx.

The fix is rarely a magic setting. It is knowing what each metric starts and stops on, which hops no metric covers, and how to put the pieces back together into one per-turn waterfall. This guide does that from the source code. Every field, default and log line below was checked against `livekit-agents` 1.8.4, the current release on PyPI, plus the LiveKit docs as of October 2026.

If you want a framework comparison, read LiveKit vs Pipecat latency. This post is for teams already on LiveKit who need to find where their milliseconds went.

1.0 s
Turn detector prediction timeout before the turn commits anyway (livekit-agents 1.8.4)
100 ms
Event-loop block that triggers a LiveKit warning span (livekit-agents 1.8.4)
505 ms
Average latency cut from speculative LLM and TTS, for 28.4% more compute (Udupa et al., Interspeech 2026)
150 ms
One-way delay ITU-T G.114 calls acceptable for most voice applications

The per-turn latency chain inside LiveKit Agents

A turn in a cascaded LiveKit agent passes through nine hops. Five of them are timed by the framework. Four are not. Here is each hop, with the mechanism that sets its duration.

1. Uplink. Caller audio travels from the handset through the carrier, into LiveKit SIP (for phone calls) or the WebRTC client, through the SFU and into your agent server. Phone audio arrives as G.711 at 8 kHz unless your trunk supports HD voice. The agent cannot see this hop. Its clock starts when audio reaches it.

2. VAD end of speech. LiveKit's inference VAD declares end of speech after `min_silence_duration` of silence (0.25 s by default in 1.8.4). Then it back-dates the "user stopped speaking" anchor: `speech_end_time = now - silence_duration - inference_duration`. Every later metric is measured from that back-dated anchor, so the VAD's silence window is already inside them.

3. STT final transcript. A turn cannot close without a final transcript. LiveKit's short-utterances guide states that interim transcripts never close a turn. The Deepgram plugin defaults `endpointing_ms` to 25, so finals arrive quickly and LiveKit's own turn logic decides the rest. The gap is `transcription_delay`.

4. Turn detector and endpointing wait. The turn detector predicts whether the user is done. The session then waits `min_delay` (likely done) or `max_delay` (likely not done), counted from the back-dated anchor. With the streaming turn detector, defaults are 0.3 s and 2.5 s. Without it, 0.5 s and 3.0 s. If the detector does not answer within 1.0 s (`DEFAULT_PREDICTION_TIMEOUT` in the source), the turn commits without a prediction and logs `eot prediction timed out, committing without a prediction`.

5. Your hook. `on_user_turn_completed` runs. RAG lookups usually live here. Its duration is `on_user_turn_completed_delay`, and it sits on the critical path every turn.

6. LLM first sentence. The LLM streams tokens. `llm_node_ttft` times the first token. A newer field, `llm_node_ttfs`, times LLM start until the first sentence reaches the TTS provider. It is in the 1.8.4 `MetricsReport` source but not yet in the docs table.

7. TTS first byte. `tts_node_ttfb` times the first audio chunk after text was sent to the provider.

8. Frame push. The first audio frame is pushed to the room track. That instant is `started_speaking_at`. The source notes that for the default room output, `playback_latency` is "self-reported when the frame is pushed to the track, so it doesn't account for network delivery to the client."

9. Downlink and playout. The frame crosses the SFU, the SIP bridge, the carrier, a jitter buffer and the handset speaker. Again, invisible to the agent.

Waterfall of one LiveKit voice turn from caller stops speaking to caller hears reply, showing the spans e2e_latency, end_of_turn_delay, llm_node_ttfs and tts_node_ttfb cover, plus network blind spots

`e2e_latency` is computed as `started_speaking_at - stopped_speaking_at`. Both timestamps are on the agent's clock, so it covers hops 2 through 8 only. The two network legs are excluded by construction.

The end-of-turn wait is a max, not a sum

The single most useful mental model for hop 4 is a max function. Approximately:

`end_of_turn_delay ≈ max(transcription_delay, endpointing_delay, VAD silence + detector inference)`

Each term starts at the same back-dated anchor. So a 250 ms VAD window and a 300 ms `min_delay` do not add to 550 ms; they overlap. But a late STT final pushes the whole turn out, no matter how low you set `min_delay`. That is why `end_of_turn_delay` is never smaller than `transcription_delay`. It is also why raising Deepgram's own endpointing to the 300 ms many guides recommend quietly raises your floor. Our stack guide covers the related trap where STT turn detection stacks `min_delay` on top of the provider's endpointing.

What each LiveKit latency metric measures and what it misses

LiveKit exposes metrics on four surfaces: per-turn `ChatMessage.metrics`, per-plugin `metrics_collected` events, live `session.usage`, and the end-of-session report. The data hooks docs mark the session-level `metrics_collected` event as deprecated. Per-plugin events are not.

The table below is organized by clock: where each timer starts, where it stops, and what it cannot see.

FieldSurfaceClock startsClock stopsBlind to
`transcription_delay`User messageBack-dated VAD end of speechLast final transcriptUplink; clamped at 0
`end_of_turn_delay`User messageBack-dated VAD end of speechTurn committedUplink
`on_user_turn_completed_delay`User messageHook calledHook returnsNothing; it is all yours
`llm_node_ttft`Assistant messageLLM generation startFirst tokenPreemptive head start
`llm_node_ttfs`Assistant messageLLM generation startFirst sentence sent to TTSCustom `tts_node` (not reported)
`llm_node_tps`Assistant messageFirst text chunkLast text chunkSingle-chunk replies
`tts_node_ttfb`Assistant messageText first sent to providerFirst audio chunkTokenizer buffering, if custom node
`e2e_latency`Assistant messageBack-dated VAD end of speechFirst frame pushed to trackBoth network legs, jitter buffer, playout
`EOUMetrics.end_of_utterance_delay`Session eventSame as `end_of_turn_delay`SameWrites 0 when the anchor is invalid
`LLMMetrics.ttft`Per-plugin eventRequest startFirst tokenReturns -1 when no tokens
`TTSMetrics.ttfb`Per-plugin eventText sentFirst audioReturns -1 when no audio

Three ways the numbers lie

Sentinels and zeros pull averages down. In 1.8.4, `_compute_end_of_turn_metrics` returns `None` when the "stopped speaking" anchor predates the turn's start. That guard exists because of issue #6093, where stale anchors produced 200-second delays. But the `EOUMetrics` event is built with `end_of_turn_delay or 0.0`, and the OpenTelemetry span uses `or 0` too. A turn the framework could not measure becomes a perfect 0 ms turn on any dashboard built on those surfaces. `LLMMetrics.ttft` and `TTSMetrics.ttfb` use -1 for "nothing generated." Average those naively and your mean improves when things break.

Preemptive generation breaks the sum. LiveKit's docs approximate total latency as `eou.end_of_utterance_delay + llm.ttft + tts.ttfb`. That holds only when the LLM starts after the turn commits. Preemptive generation, which is on by default in 1.8.4, starts the LLM earlier. Then `llm_node_ttft` overlaps the end-of-turn wait, and the sum overstates the real gap.

Missing turns hide the worst turns. A user turn that never produced agent audio (interrupted, skipped, or blocked by `min_words`) gets no `e2e_latency`. If your slowest turns are the ones where callers gave up and spoke again, they vanish from the percentile.

The reconciliation check: residual against e2e_latency

Here is an original diagnostic that uses only fields LiveKit already records. For each turn, compute:

`residual = e2e_latency - (end_of_turn_delay + on_user_turn_completed_delay + llm_node_ttfs + tts_node_ttfb)`

Use `llm_node_ttft` in place of `llm_node_ttfs` when the latter is absent. Read the sign:

  • Residual near zero (within about 50 ms): the waterfall is complete. Fix the biggest component.
  • Residual strongly negative: preemptive generation is working. The magnitude is roughly the head start it bought you. A turn with a large negative residual and a short `end_of_turn_delay` is preemption done right.
  • Residual strongly positive: time is being lost between stages. The usual suspects are a tool call (the first LLM pass returned a function call, so speech waited for the tool plus a second LLM pass), speech scheduling behind an uninterruptible utterance, or a blocked event loop.

The positive case is the one teams miss. In the 3-4 s community thread above, the poster's component metrics summed to under a second. A residual computed per turn would have shown immediately whether the gap was inside the agent (positive residual) or outside it (residual near zero, but callers still waiting). The most specific reply in that thread pointed downstream of the agent, at client playout, though the poster never confirmed a root cause.

Code: per-turn waterfalls with p50 and p95

The collector below writes one JSON line per answered user turn, pairing each user message with the assistant message that answered it. It uses `conversation_item_added` and `ChatMessage.metrics`, which are not deprecated. It is illustrative and targets `livekit-agents` 1.8.x; adapt the plumbing to your entrypoint.

# latency_waterfall.py - illustrative, livekit-agents 1.8.x
import json
from livekit.agents import AgentSession, ConversationItemAddedEvent
from livekit.agents.llm import ChatMessage

USER_FIELDS = ("transcription_delay", "end_of_turn_delay", "on_user_turn_completed_delay")
AGENT_FIELDS = ("llm_node_ttft", "llm_node_ttfs", "llm_node_tps",
                "tts_node_ttfb", "playback_latency", "e2e_latency")

def attach_waterfall_logger(session: AgentSession, call_id: str, path: str) -> None:
    state = {"user": None, "turn": 0, "unanswered": 0}
    out = open(path, "a", buffering=1)

    @session.on("conversation_item_added")
    def _on_item(ev: ConversationItemAddedEvent) -> None:
        item = ev.item
        if not isinstance(item, ChatMessage):
            return
        m = dict(item.metrics or {})
        if item.role == "user":
            if state["user"] is not None:
                state["unanswered"] += 1      # caller spoke again before a reply
            state["user"] = {k: m.get(k) for k in USER_FIELDS}
            return
        if item.role != "assistant" or state["user"] is None:
            return                             # greeting or say(): no user turn
        state["turn"] += 1
        row = {"call_id": call_id, "turn": state["turn"],
               "unanswered_before": state["unanswered"],
               **state["user"], **{k: m.get(k) for k in AGENT_FIELDS}}
        out.write(json.dumps(row) + "\n")
        state["user"], state["unanswered"] = None, 0

Per-plugin LLM events add two signals the per-turn report lacks: cache hits and cancelled generations. `LLMMetrics` carries `prompt_tokens`, `prompt_cached_tokens` and a `cancelled` flag. Discarded preemptive generations show up as cancelled requests, which is how you measure what preemption costs.

# attach to the LLM instance you construct, e.g. llm = openai.LLM(...)
def attach_llm_logger(llm, call_id: str, path: str) -> None:
    out = open(path, "a", buffering=1)
    def _on_metrics(m) -> None:                # m is an LLMMetrics
        out.write(json.dumps({
            "call_id": call_id, "speech_id": m.speech_id, "ttft": m.ttft,
            "cancelled": m.cancelled, "prompt_tokens": m.prompt_tokens,
            "cached": m.prompt_cached_tokens}) + "\n")
    llm.on("metrics_collected", _on_metrics)

Then aggregate offline. The script drops `None`, zeros and sentinels from percentiles but counts them, because coverage is a metric in its own right.

# waterfall_report.py - illustrative
import json, numpy as np

rows = [json.loads(l) for l in open("turns.jsonl")]
STAGES = ["end_of_turn_delay", "on_user_turn_completed_delay",
          "llm_node_ttfs", "tts_node_ttfb", "e2e_latency"]
MAX_DELAY = 2.5   # your endpointing max_delay

def clean(vals):
    return np.array([v for v in vals if isinstance(v, (int, float)) and v > 0])

for s in STAGES:
    raw = [r.get(s) for r in rows]
    v = clean(raw)
    cov = len(v) / max(len(raw), 1)
    if len(v):
        p50, p95 = np.percentile(v, [50, 95]) * 1000
        print(f"{s:30s} p50={p50:6.0f} ms  p95={p95:6.0f} ms  coverage={cov:.1%}")

eot = clean([r.get("end_of_turn_delay") for r in rows])
print(f"max-delay rate: {(eot >= 0.8 * MAX_DELAY).mean():.1%}")

resid = []
for r in rows:
    first = r.get("llm_node_ttfs") or r.get("llm_node_ttft")
    parts = [r.get("end_of_turn_delay"), r.get("on_user_turn_completed_delay") or 0,
             first, r.get("tts_node_ttfb")]
    if r.get("e2e_latency") and all(isinstance(p, (int, float)) for p in parts):
        resid.append(r["e2e_latency"] - sum(parts))
resid = np.array(resid) * 1000
print("residual ms p05/p50/p95:", np.percentile(resid, [5, 50, 95]).round())
print("turns > 4 s:", (clean([r.get("e2e_latency") for r in rows]) > 4).mean())

The max-delay rate deserves its own alert. End-of-turn delays are bimodal: they cluster near `min_delay` when the detector is confident and near `max_delay` when it is not. A healthy p50 can sit on top of a p95 made entirely of "detector unsure" turns.

Measure true end to end from the caller side

`e2e_latency` stops when a frame is pushed to the track. Callers start listening for it much later. To measure what they hear, you need audio from beyond the agent.

The cleanest method is a triangulation across three vantage points:

1. Agent metrics. `e2e_latency` per turn, from the collector above.

2. SFU-side recording. A LiveKit Egress track recording captures audio as the room sees it, after the agent has published and before the SIP bridge and carrier.

3. Far-end recording. A recording made at the caller's handset, or at the carrier on a two-channel (stereo) call recording, with caller and agent on separate channels.

Measure the gap from the end of caller speech to the onset of agent speech in recordings 2 and 3, then subtract. Recording 2 minus agent metrics is the agent-to-room cost. Recording 3 minus recording 2 is the SIP bridge, carrier, jitter buffer and playout.

That is exactly the method the reporter of GitHub issue #3685 used informally: comparing an Egress recording against a microphone recording on the phone. The response gap was far larger on the phone side, which pointed at the telephony leg, not the pipeline. A LiveKit team member replied in the thread that the LiveKit-to-Twilio leg "should be around 100-150 ms," with the rest from Twilio to the end user over the PSTN.

For onset detection, use hysteresis rather than a single energy threshold, and require a minimum duration so a click or a breath does not count as speech. This sketch returns per-turn gaps you can line up with the collector's rows by order.

# turn_gaps.py - illustrative onset detection with hysteresis
import numpy as np, soundfile as sf

def active(ch, sr, on=0.03, off=0.015, min_ms=120, hop_ms=10):
    hop = int(sr * hop_ms / 1000)
    n = len(ch) // hop
    rms = np.sqrt((ch[: n * hop].reshape(n, hop) ** 2).mean(axis=1))
    state, run, out = False, 0, np.zeros(n, bool)
    for i, e in enumerate(rms):
        if not state and e > on:
            run += 1
            if run * hop_ms >= min_ms:
                state, run = True, 0
                out[max(0, i - min_ms // hop_ms + 1) : i + 1] = True
        elif state and e < off:
            state = False
        else:
            run = 0 if not state else run
        out[i] |= state
    return out

audio, sr = sf.read("call_stereo.wav")      # ch0 caller, ch1 agent
caller, agent = active(audio[:, 0], sr), active(audio[:, 1], sr)
gaps, i = [], 1
while i < len(caller):
    if caller[i - 1] and not caller[i]:      # caller just stopped
        j = i
        while j < len(agent) and not agent[j] and not caller[j]:
            j += 1
        if j < len(agent) and agent[j]:
            gaps.append((j - i) * 10)        # ms of silence the caller heard
        i = j
    i += 1
print("caller-heard gap p50/p95 ms:", np.percentile(gaps, [50, 95]))

Two cautions. Mid-sentence pauses longer than your endpointing will register as false turn ends, so set `min_ms` and review a sample by ear. And on a carrier recording the agent channel already includes the downlink, but the caller channel has not yet crossed the uplink, so the measured gap understates what the caller heard by roughly one uplink leg. Note which recording you used whenever you report a number.

Region placement and RTT math

Distance is the latency you cannot tune away with settings. LiveKit's own latency guide ranks agent-model co-location as the highest-impact fix.

The physics floor is simple. Light in fiber covers about 200,000 km per second, so the minimum round-trip time is:

`RTT_min (ms) = 2 x distance_km / 200` which is about 1 ms per 100 km.

Real routes run longer than the great-circle line. As a working assumption, take 1.5 times the floor plus 10 ms. Then count how many serial round trips a turn makes to each remote party. On warm connections:

Remote callSerial round trips per turnNote
Streaming STT finalAbout 0.5Audio streams continuously; the final travels one way
Cloud turn detector inferenceAbout 1Only if the model runs remotely
LLM request to first token1Plus 2 more if the connection is cold (TCP and TLS 1.3)
TTS text to first audio1Websocket already open
Each tool call1 LLM + 1 APIThe extra LLM pass that speaks the tool result

Now plug in three placements. All numbers are worked examples, not measurements.

  • A. Co-located. Agent and model endpoints in one metro, RTT about 2 ms. A no-tool turn spends under 10 ms on distance.
  • B. Agent in Virginia, LLM and TTS endpoints in Oregon. About 3,700 km, floor 37 ms, assumed real RTT 65 ms. A no-tool turn pays about 2.5 RTTs (STT, LLM, TTS), around 160 ms. One tool call adds about two more RTTs (the second LLM pass and the API call, if the API is equally far), for roughly 290 ms of pure distance.
  • C. Agent in Mumbai, models in US East. About 13,000 km, floor 130 ms, assumed real RTT 205 ms. A no-tool turn pays around 510 ms in distance alone, before any model computes anything. One tool call pushes that past 900 ms.
Map of the LiveKit call path from caller through carrier edge, SIP region, SFU and agent server to model providers, with serial round trips and three placement scenarios

LiveKit placement rules that surprise people

  • Inbound SIP enters near your trunk, not your caller. The region pinning docs say incoming calls route "to the region closest to the SIP trunking provider's endpoint." With Twilio, that endpoint is the edge you configure. Twilio's edge list offers `ashburn`, `umatilla` and others, and states that the `roaming` low-latency edge "isn't available for SIP Domains or Elastic SIP Trunking." You pick the edge, so you pick the path.
  • Outbound calls start where your API call runs. The same page: "Outgoing calls originate from the same region where the `CreateSIPParticipant` API call is made." A backend in Europe dialing US numbers sends media through Europe unless you set `destination_country` on the trunk. The issue #3685 thread shows this pattern: calls routed through one region while the carrier used another.
  • Agent regions are fixed and can spill. The agent deployment docs say a deployment's region "can't be changed after creation," and that if agents are at capacity, "users may connect to an agent in a different region." A capacity shortfall shows up as a latency tail, not an error.
  • Region-pinned SIP endpoints exist. The format is `{sip_subdomain}.{region}.sip.livekit.cloud`, with regions `aus`, `eu`, `india`, `japan`, `sa`, `uk` and `us`.

For context on the phone leg itself, ITU-T G.114 treats 0 to 150 ms of one-way transmission delay as acceptable for most voice applications and 400 ms as the planning limit. Two network legs at 150 ms each already consume 300 ms. Human conversation leaves little room for that: Stivers et al. (PNAS, 2009) found that in all 10 languages studied, the most common turn transition fell between 0 and 200 ms. Our SIP vs WebRTC guide and telephony comparison go deeper on the transport side.

Know your caller-heard p95, not just e2e_latency
Evalgent runs independent test calls through your real SIP path and reports caller-heard latency per turn next to your LiveKit metrics, so you can see which segment regressed.
Book a demo

A worked waterfall with p50 and p95 budgets

Here is a budget for a US inbound phone agent with co-located models. Every number is an illustrative target, not a measurement. Replace each with your own percentiles from the report script.

SegmentMetricp50 budget (ms)p95 budget (ms)Usually blown by
Uplink: handset, carrier, SIP, SFURecording only90160Wrong trunk edge, cross-region media
End-of-turn wait (includes VAD and STT final)`end_of_turn_delay`4201,250Detector unsure, so `max_delay` wins
Your hook`on_user_turn_completed_delay`15120Synchronous RAG or database call
LLM to first sentence`llm_node_ttfs`4501,100Long prompt, cache miss, tool call
TTS first audio`tts_node_ttfb`160380Cold websocket, provider region
Scheduling residualComputed2090Event-loop stalls
Downlink, jitter buffer, playoutRecording only110200Carrier routing, client playout
Agent-side `e2e_latency`Sum of rows 2 to 61,065see below
Caller-heardRecording1,265see below

At the median the arithmetic is clean: 420 + 15 + 450 + 160 + 20 = 1,065 ms agent-side, plus 200 ms of network for 1,265 ms caller-heard.

At p95 it is not. Summing the p95 column gives 3,300 ms caller-heard, but that number describes a turn where every stage hits its tail at once, which almost never happens. Do not budget p95 by summing p95s. Budget the p95 of the total directly, and use the per-stage p95s to find which stage owns the tail.

Dean and Barroso's "The Tail at Scale" (CACM, 2013) explains why the tail still matters so much. Their example: a server that is slow 1% of the time makes 63% of requests slow once a request touches 100 such servers. A voice turn is a short serial chain, not a 100-way fan-out, but the same arithmetic applies. If each of five stages independently lands in its own slowest 5% on a given turn, the chance a turn hits at least one tail is:

`1 - 0.95^5 = 22.6%`

Nearly one turn in four contains at least one stage tail. In practice that means the p95 turn is usually one stage misbehaving, not all of them. Your waterfall should tell you which one.

The same paper offers a fix that maps directly onto voice: hedged requests. Send a backup request only after the first has been outstanding longer than its p95, and cancel the loser. In their BigTable benchmark, hedging after 10 ms cut 99.9th-percentile latency from 1,800 ms to 74 ms with 2% more requests. LiveKit's `FallbackAdapter` is a simpler cousin: its `attempt_timeout` defaults to 5.0 s, which is too slow for a conversational turn.

Fixes by segment

Endpointing and the turn detector

If `end_of_turn_delay` dominates and the max-delay rate is high, the detector is unsure. Short replies like "yes" and "okay" are the classic case. In a community thread on Agents 1.6.4, every short reply showed `eou_delay = 1.00s`, exactly the configured `max_delay`. The poster reported that switching to `dynamic` endpointing worked better for short utterances.

Dynamic mode adapts the wait using an exponential moving average of the caller's pauses (`alpha` defaults to 0.9). You can also update endpointing at runtime for a known yes/no stretch of the conversation. Do not just crush `max_delay`. LiveKit's short-utterances guide warns that overly aggressive endpointing ends turns early and can attach a final transcript to the wrong turn. Our endpointing comparison covers the trade-off in detail.

Preemptive generation, and the RAG pattern that silently disables it

In 1.8.4, `PreemptiveGenerationOptions` defaults to `enabled: True`, `preemptive_tts: False`, `max_speech_duration: 10.0` and `max_retries: 3`. It fires on each new final or preflight transcript, before the turn is committed, and skips utterances longer than 10 seconds.

Here is the gotcha that is not in the docs. When the turn commits, the framework reuses the preemptive reply only if the transcript, the chat context, the tools and the tool choice are all unchanged after `on_user_turn_completed` runs. The standard RAG pattern adds retrieved context to the turn's chat context inside that hook. That changes the chat context, so every preemptive generation is thrown away. You pay full LLM latency plus the wasted tokens. The framework logs this at warning level. Grep for the stable prefix, because the full message wraps the hook name in backticks:

`preemptive generation invalidated after`

The full warning ends "because the transcript, chat context, tools, or tool choice changed."

When reuse works, the framework logs `using preemptive generation` at debug level with a `preemptive_lead_time` field. Count both lines per session and you have a preemptive hit rate. To keep RAG and preemption together, retrieve during the user's speech (keyed on the interim or preflight transcript) and put the results somewhere that does not mutate the turn context at commit time, or accept the trade-off consciously.

Research shows what a good speculative system buys. Udupa et al. (Interspeech 2026) built a model that forecasts the end of turn up to 2.56 s ahead and starts the LLM and TTS speculatively. Integrated into the Unmute framework, it cut average latency by 505 ms at the cost of 28.4% more speculative computation. That is the right frame for LiveKit too: track latency saved and tokens wasted together, using the negative residual and the `cancelled` flag on `LLMMetrics`.

LLM time to first token and first sentence

TTFT grows with prompt length because the model must process every input token before emitting one. Two levers matter most in voice:

  • Keep the prefix cacheable. OpenAI's prompt caching docs say cache reuse "requires the entire rendered prefix to match," with a 1,024-token minimum for GPT-5.6 and later. Put static instructions and tool definitions first. Put the caller's name, today's date and per-call variables last. Watch `prompt_cached_tokens / prompt_tokens` from `LLMMetrics`; a ratio near zero on turn three means something at the top of your prompt changes every request.
  • Count tool steps. Each tool call means a second LLM pass before speech. LiveKit's guide recommends limiting `max_tool_steps` and playing a thinking sound. Our post on agents that go silent covers what callers do during that silence.

Then check the retry defaults, because they shape your p99. `APIConnectOptions` defaults to `max_retry=3`, `retry_interval=2.0` and `timeout=10.0`. The first retry waits 0.1 s, later ones 2.0 s. A provider that hangs on connect can hold a turn for tens of seconds before the session sees an error. The OpenAI plugin's HTTP client separately uses a 15 s connect timeout and a 5 s read timeout. For a conversation, set tighter values through `SessionConnectOptions(llm_conn_options=...)` and fail over to a second model faster.

TTS streaming and sentence chunking

`llm_node_ttfs` exists because the first audio cannot start until a first sentence reaches the TTS. LiveKit's default blingfire sentence tokenizer uses `min_sentence_len=20`: spans shorter than 20 characters merge into the next sentence. So a reply that opens with "Sure!" does not ship "Sure!" early. It waits for the next sentence to finish. A first clause of 20 to 60 characters ships fastest. You can steer that with one line in your prompt.

Plugin choices change this too:

  • Non-streaming TTS runs through `StreamAdapter`, which synthesizes one whole sentence at a time, so its first-byte time includes synthesizing that full sentence. In one community thread, a poster stuck at 3 seconds per turn reported about 1 second after switching TTS providers, while one other TTS model stayed at three seconds. The thread does not pin down the cause, but provider choice was the variable that moved.
  • ElevenLabs: the plugin uses `auto_mode` with a sentence tokenizer by default. Passing `chunk_length_schedule` turns `auto_mode` off, and the provider then buffers text per that schedule before generating. The plugin documents `[120, 160, 250, 290]` characters as the schedule's defaults, so the first audio waits for about 120 characters of LLM output.
  • Custom `tts_node`: `tts_node_ttfb` falls back to timing from the first input token, so tokenizer buffering is counted as TTFB and `llm_node_ttfs` is not reported.

See our low-latency TTS guide and time to first audio explainer for provider-side detail.

Cold starts and prewarm

The first turn of a call pays costs later turns do not. In 1.8.4, `AgentServer` keeps `num_idle_processes` warm: 0 in dev mode, `ceil(cpu_count)` in production. Load VAD weights in a setup function (`server.setup_fnc = prewarm`) so jobs do not load them on pickup. The LLM's `prewarm()` resolves DNS and opens the TLS connection, and the source says it "is called automatically when an `AgentSession` is constructed and when an agent activity starts." LiveKit's guide adds that Build-plan projects get several-second cold starts once all sessions end. Report first-turn latency separately so it does not pollute steady-state percentiles.

Event-loop stalls

A voice pipeline runs VAD, playout and every provider stream on one asyncio loop. Any synchronous call, such as a blocking HTTP client inside a tool, stalls all of them. Version 1.8.4 ships a loop monitor that logs `event loop blocked for {N}ms` at a 100 ms warning threshold and 500 ms error threshold. You can change them with `LIVEKIT_AGENTS_LOOP_BLOCK_WARN_MS` and `LIVEKIT_AGENTS_LOOP_BLOCK_ERROR_MS`. The session span carries `lk.blocking.count`, `lk.blocking.total_duration` and `lk.blocking.max_duration`, and each stall span includes a sampled stack. If `lk.blocking.cpu_time` is near zero, the process was not running at all, which points at CPU quotas or burstable instances rather than your code. Load-dependent stalls show up only under concurrency, so pair this with stress testing.

Symptom to segment to fix: a LiveKit latency runbook

Decision tree that routes a LiveKit latency symptom to the responsible segment using e2e_latency, the residual, the max-delay rate and recordings, ending in a specific fix for each branch
SymptomCheck firstLikely segmentFix
Fast in browser, slow on phoneRecording 3 minus recording 2SIP and PSTN legsPin trunk edge and LiveKit SIP region near callers; set `destination_country` for outbound
Metrics sum under 1 s, callers wait 3 sResidual near zero, recordingsDownlink or client playoutTest with a known-good client; inspect custom player buffering
Short replies wait exactly `max_delay`Max-delay rateTurn detector unsureDynamic endpointing; runtime endpointing for yes/no prompts
`end_of_turn_delay` never below about 0.5 s`transcription_delay`STT final arrives lateLower provider endpointing; in STT mode set `min_delay` to 0
Strongly positive residual on some turnsTool calls per turnTool round tripsCap `max_tool_steps`; consolidate APIs; thinking sound
No negative residuals at allGrep "preemptive generation invalidated"Preemption discardedMove RAG out of the chat-context mutation in the hook
`llm_node_ttft` creeps up over a call`prompt_cached_tokens` ratioPrompt growth, cache missStatic prefix first; trim or summarize history
`llm_node_ttfs` far above `llm_node_ttft`First sentence lengthSentence tokenizerPrompt a 20 to 60 character opening clause
High `tts_node_ttfb` on first turn onlyTurn indexCold connectionRely on prewarm; keep a session-level TTS instance
Random 1 to 10 s spikes`event loop blocked` logs, retry logsLoop stall or provider retryMove sync work off-loop; tighten `SessionConnectOptions`
Latency rises with call volumep95 by concurrencyCPU or region spillCompute-optimized instances; capacity per region
First turn of every call slowFirst-turn p50Process or model cold startIdle processes; prewarm VAD; paid plan

How to test LiveKit latency before every release

A latency fix that is not measured the same way twice is a guess. This protocol produces numbers you can compare across builds.

1. Pin the environment. Record the `livekit-agents` version, turn detector version, endpointing mode and values, preemptive options, model names and regions. A change to any of them is a new baseline. Our LLM update regression guide applies equally to framework upgrades.

2. Build a scripted turn set. Use at least 40 calls of 10 to 12 turns each, so 400 or more answered turns. Mix short replies ("yes", a ZIP code), long multi-clause requests, turns that trigger a tool, and turns that do not. Use synthetic callers so the same script runs every time.

3. Run each set under four conditions. WebRTC and SIP; single call and your target concurrency. Keep first turns in a separate bucket.

4. Collect three vantage points. The waterfall JSONL, an Egress recording, and a two-channel far-end recording for at least a subset of calls.

5. Compute the report. Per-stage p50 and p95, metric coverage, max-delay rate, residual distribution, preemptive hit rate, cancelled-token share, loop stalls, and caller-heard p50 and p95 from audio.

6. Apply pass thresholds. These are starting points we suggest, not an industry standard. Tune them to your use case.

CheckPass threshold
Agent-side `e2e_latency` p50 / p951,000 ms / 1,800 ms or better
Caller-heard p50 / p95 (SIP)1,300 ms / 2,200 ms or better
Turns above 4 s caller-heard1% or fewer
Metric coverage on answered turns97% or more
Max-delay rate10% or less
Preemptive hit rate (if enabled)60% or more
Loop stalls above 500 ms0 per 100 turns

The 4-second line is anchored in research: Maslych et al. (CUI 2025) found that response latency above 4 seconds degraded quality of experience in conversations with LLM-powered agents, while natural conversational fillers improved perceived response time.

7. Size the sample for the tail. A p95 from 100 turns is close to noise. Using the binomial approximation, the rank of the true p95 falls within `n x 0.95 ± 1.96 x sqrt(n x 0.95 x 0.05)`. For n = 100 that is ranks 91 to 99, so your "p95" could be anywhere from the 91st to the 99th percentile. For n = 400 it is ranks 372 to 389, roughly the 93rd to the 97th percentile. Turns within one call are correlated, so spread them across many calls.

8. Compare builds on distributions. Bootstrap the difference in p95 between the old and new build (resample whole calls, 2,000 iterations) and ship only if the 95% interval excludes a regression larger than your tolerance, say 150 ms.

Where independent evaluation fits

Everything above works with your own tooling, and you should run it. The gap is coverage. Engineering teams test the paths they wrote, from the network they sit on, on the build they just changed. Caller-heard latency on real carrier routes, at production concurrency, across every prompt and model update, is the part that slips.

That is where Evalgent fits. We run independent, scripted test calls through your actual phone numbers and SIP trunks, score caller-heard latency per turn alongside conversational quality, and flag regressions between releases. Teams use it for a pre-launch audit, for comparing STT, LLM and TTS vendors on the same scripted calls, and as a release gate. It does not replace your LiveKit metrics. It tells you whether the number your callers experience matches the one in your logs. For what to capture on every call, see what to log on voice agent calls, and for the broader test plan, our LiveKit testing guide.

Frequently asked questions

Why do LiveKit metrics say 900 ms when callers hear 2 seconds?

`e2e_latency` runs from the back-dated end of user speech to the moment the agent pushes its first audio frame to the room. It never sees the uplink, the downlink, the SIP bridge, the carrier, the jitter buffer or the handset. Measure from a far-end recording and subtract to find the gap.

What does end_of_turn_delay include?

It runs from the back-dated VAD end of speech to the moment the turn commits. That window already contains the VAD silence, the wait for the STT final transcript and the endpointing delay. Those overlap rather than add, so it behaves roughly like the maximum of the three.

Why is my LiveKit EOU delay always exactly max_delay?

The turn detector judged the turn likely incomplete, so the session waited the full `max_delay`. Short replies like "yes" trigger this often. Try dynamic endpointing, update endpointing at runtime for yes/no stretches, and track the share of turns that hit 80% or more of `max_delay`.

Is preemptive generation on by default in LiveKit?

In `livekit-agents` 1.8.4, yes. `PreemptiveGenerationOptions` defaults to enabled for the LLM, with `preemptive_tts` off, a 10-second speech cap and three attempts per turn. It is discarded if `on_user_turn_completed` changes the transcript, chat context, tools or tool choice.

What is llm_node_ttfs in LiveKit?

It is the time from LLM generation start until the first sentence reaches the TTS provider, as segmented by that TTS. It appears in the 1.8.4 `MetricsReport` source alongside `llm_node_tps`. It captures sentence-tokenizer buffering that `llm_node_ttft` misses, and it is not reported with a custom `tts_node`.

Does LiveKit SIP add latency?

LiveKit's guide says SIP's impact is generally limited to round-trip time plus small jitter and transcoding allowances, unless routing adds distance. Inbound calls enter near your trunk provider's endpoint and outbound calls start where you call `CreateSIPParticipant`, so trunk edge and API region choices decide that distance.

How many test turns do I need to trust a p95?

At least 400 answered turns spread across 40 or more calls. With 400 samples the measured p95 sits roughly between the true 93rd and 97th percentiles. With 100 samples it could be anywhere from the 91st to the 99th, which is too wide to catch a regression.

Should I average LiveKit latency metrics?

No. Some surfaces record unmeasurable turns as 0 and failed generations as -1, which pulls averages down exactly when things break. Report p50 and p95, exclude sentinels from percentiles, and track metric coverage as its own number.

The bottom line

A LiveKit latency issue is almost always one segment, and you can find it by rebuilding each turn's waterfall from `ChatMessage.metrics`, checking the residual against `e2e_latency`, and comparing against what a recording says the caller heard. Fix the segment that owns your p95, then lock the result in with a fixed test protocol so the next prompt, model or framework upgrade cannot quietly undo it.

Related Articles