Evalgent
Back to Blog
Voice AI Testing

What to Log on Every Voice Agent Call: The Complete Schema

Deepesh Jayal
24 min read
What to Log on Every Voice Agent Call: The Complete Schema

A caller says the agent "just went silent." You open your logs. You find a transcript and a timestamp. You do not find the model snapshot, the endpointing setting, the tool response, or the moment the TTS stream stalled. The ticket closes as "could not reproduce."

That gap is the whole problem. Most teams log what their platform hands them and nothing more. Then they try to test, evaluate and debug on top of it. It does not work.

This post is the logging standard we wish every team started with. It gives you the field catalog as tables. It includes a JSON record for one call and one turn. It shows Python that emits structured logs and OpenTelemetry spans, plus the real hook names in LiveKit and Pipecat. It ends with ten SQL queries your logs must answer. Every provider field name below was checked against the vendor's own docs in September 2026. Where we simplified, we say so.

Voice agent call log: a structured, joinable record of everything that happened on one call. It covers configuration, timing, recognition, reasoning, tools, speech output, outcome and cost. It is keyed so any turn can be replayed and scored later.

Why call logging comes before testing, evaluation and debugging

You cannot replay what you did not capture. You cannot score a turn you cannot find. You cannot explain a failure without the inputs that caused it.

Testing needs real inputs. The best regression suites are built from production calls. That requires the caller audio, the transcript the agent actually heard, and the config it ran with.

Evaluation needs ground truth and context. A judge that sees only the transcript misses the 2.4 seconds of dead air. It misses the tool that returned an empty body. It misses the barge-in the agent ignored.

Debugging needs a timeline. Voice failures are timing failures more often than logic failures. The words were right, but they arrived late, overlapped the caller, or never played.

We use one test for every logging setup we review.

The five-minute test: pick a random call ID from yesterday. Can one engineer rebuild it in under five minutes? That means what the caller said, what the agent heard, and what it decided. It means which tools ran, what they returned, what the agent said, and how long each step took. If not, your logging is not done.

Most teams fail on three points. They lack version info, so they cannot tie a change to a regression. They lack per-turn timing, so latency is one averaged number. They lack provider request IDs, so vendor tickets stall. For a broader framing, see testing vs monitoring vs observability.

The four layers of a voice agent call log

A voice call has natural grain. Log at each grain and join them with stable IDs.

The four layers of a voice agent call log: the call envelope, per-turn records, a fine-grained event timeline, and artifacts like audio and transcripts, joined by call, turn and provider IDs
LayerGrainRows per callWhat it answersTypical store
Call (envelope)One per call1Who, when, which versions, outcome, costWarehouse table
TurnOne per user-agent exchange5 to 60Latency, what was heard, what was said, toolsWarehouse table
EventOne per state change200 to 5,000Exact ordering, gaps, overlaps, retriesLog store or trace backend
ArtifactOne per file2 to 6Audio, transcripts, provider reportsObject storage

Row counts are illustrative for a three-to-ten-minute support call.

Call envelope: the single summary row for a call. It holds identity, versions, telephony facts, outcome and cost totals.

Turn record: one row per exchange. It starts when the user stops speaking and ends when the agent finishes, is interrupted, or yields.

Event timeline: append-only, timestamped events. Examples are VAD start, STT final, LLM first token, tool start, TTS first byte and playback start.

Artifact: a file referenced by URI, never inlined. Examples are dual-channel audio, raw and redacted transcripts, and provider session reports.

The IDs that join everything

Joins break more investigations than missing data does. Generate your own call_id at the first touchpoint. Then attach every provider ID you can get.

IDWho creates itWhere you get itJoin to
`call_id`YouGenerate at call start (UUIDv7 sorts by time)Everything
`turn_id`You`call_id` plus a turn indexEvents, tool calls, eval results
`trace_id` / `span_id`OpenTelemetry SDKActive span contextCall and turn rows
`CallSid`TwilioStatus callbacks and the Media Streams `start.callSid`Call envelope
`StreamSid`TwilioMedia Streams `start.streamSid`Call envelope, audio events
`ParentCallSid`TwilioStatus callback on child legsTransfer legs
`request_id`Deepgram`metadata.request_id` on each `Results` messageSTT events, vendor tickets
`speech_id`LiveKit AgentsEOU, LLM and TTS metrics eventsYour `turn_id`
`conversation.id`PipecatConversation span attribute, or `conversation_id` you passYour `call_id`
`item_id`, `response_id`OpenAI RealtimeServer events such as `input_audio_buffer.speech_started` and `response.done`Turn and event rows
`call.id`VapiCall object in every server messageCall envelope
`call_id`RetellCall object in webhooks and `get-call`Call envelope
`tool_call_id`LLM provider or platformTool call payloadTool call rows

Push your call_id into provider metadata wherever possible. Twilio `` custom parameters come back in `start.customParameters`. Deepgram accepts `tag` and `extra` query parameters on the streaming request. Retell has a free-form `metadata` object on the call. That way, a vendor's own record points back at yours.

The field catalog

This is the core of the standard. Each table lists the field, a type, and why it earns its place. Names are our recommended schema. Provider field names appear in code formatting.

Identity and versioning

Version fields are what make regressions findable. Without them, "latency got worse on Tuesday" has no suspect.

FieldTypeExampleWhy it matters
`call_id`string`01J9Z...`Primary key for every join
`agent_id`string`billing-agent`Which agent handled the call
`agent_version`string`2026.09.24-3`Deploy-level version, set in CI
`prompt_id` / `prompt_version`string`billing/v41`Human-readable prompt version
`prompt_hash`string`sha256:9f2c...`Catches edits that skipped a version bump
`stt_provider` / `stt_model` / `stt_model_version`string`deepgram` / `nova-3` / from `model_info.version`STT drift and snapshot changes
`llm_provider` / `llm_model_requested` / `llm_model_served`string`openai` / `gpt-4.1` / served snapshotAliases resolve to new snapshots silently
`tts_provider` / `tts_model` / `voice_id`string`cartesia` / `sonic-2` / `a0e9...`Voice and prosody regressions
`config_hash`string`sha256:41ab...`Hash of endpointing, VAD, interruption and timeout settings
`tool_schema_version`string`tools/v12`Argument errors after a schema change
`feature_flags`map`{"preemptive_generation": true}`A/B and rollout attribution
`experiment_id` / `variant`string`exp-77` / `B`Clean comparisons
`deployment_env` / `region` / `host`string`prod` / `us-east-1` / pod nameRegion-specific latency and outages
`framework` / `framework_version`string`livekit-agents` / `1.7.x`Attribute renames and behavior changes

Log `llm_model_served` separately from what you requested. The OpenTelemetry GenAI conventions make the same split with `gen_ai.request.model` and `gen_ai.response.model`. An alias can move under you. See how LLM updates cause voice agent regressions.

Telephony and audio

FieldTypeExampleSource
`direction`enum`inbound`, `outbound-api`, `outbound-dial`Twilio `Direction`
`from_number_hash` / `to_number_hash`stringsalted hashHash, do not store raw by default
`carrier` / `trunk_id`stringcarrier name, `TK...`Your SIP provider, Twilio `trunk_sid`
`transport`enum`pstn`, `sip`, `webrtc`Platform config
`codec` / `sample_rate_hz` / `channels`string, int, int`audio/x-mulaw`, `8000`, `1`Twilio `start.mediaFormat`
`call_status_final`enum`completed`, `busy`, `failed`, `no-answer`, `canceled`Twilio `CallStatus`
`sip_response_code`int`200`, `404`, `487`Twilio `SipResponseCode`, terminal events only
`answered_by`enum`human`, `machine`, `unknown`Twilio `AnsweredBy` when AMD is on
`ring_ms` / `answer_ts` / `end_ts`int, timestamp`4200`Status callback events
`recording_ref`URI`s3://calls/raw/...`Your storage, not the vendor URL
`recording_channels`enum`dual`Twilio `RecordingChannels`
`dtmf_count`int`3`Count only; never log digits in payment flows

Twilio sends `initiated`, `ringing`, `answered` and `completed` progress events when you request them via `StatusCallbackEvent`. `CallDuration` and `SipResponseCode` only arrive on terminal events. Record dual-channel audio whenever consent allows. Mono audio makes overlap and barge-in analysis much harder.

Timing per turn

Timing is where voice agents fail most, and where logs are thinnest. Log absolute timestamps for each milestone. Derive durations later.

The timing fields to log for one voice agent turn: user speech end, endpoint decision, STT final, LLM first token, tool call, TTS first byte and first audio played, and the user-perceived latency they add up to
FieldMeaningWhere to get it
`user_speech_start_ms`VAD detected speech onsetPipecat VAD frames, OpenAI `input_audio_buffer.speech_started` (`audio_start_ms`), Deepgram `SpeechStarted`
`user_speech_end_ms`VAD detected speech endLiveKit `stopped_speaking_at`, OpenAI `input_audio_buffer.speech_stopped` (`audio_end_ms`)
`endpoint_decision_ms`Turn declared completeLiveKit `end_of_turn_delay`, Pipecat `TurnMetricsData`
`stt_final_ms`Final transcript availableLiveKit `transcription_delay`, Deepgram `speech_final: true`
`llm_request_ms`LLM request sentYour code
`llm_first_token_ms`First token receivedLiveKit `llm_node_ttft`, Pipecat `TTFBMetricsData` / `TTFATMetricsData`
`first_tool_call_ms`First tool call emittedYour tool wrapper
`tool_result_ms`Last tool result returnedYour tool wrapper
`tts_request_ms`First text sent to TTSYour code
`tts_first_byte_ms`First audio chunk from TTSLiveKit `tts_node_ttfb`, Pipecat `metrics.ttfb` on the TTS span
`first_audio_played_ms`First agent audio sent to transportLiveKit `playback_latency`, Twilio `mark` echo
`agent_speech_end_ms`Agent finished or was cut offYour transport, Twilio `mark`
`barge_in_ms`Caller started speaking over the agentVAD during agent playback

Then derive the one number callers feel.

User-perceived latency: the time from the end of the caller's speech to the first agent audio reaching the caller. In logs it is `first_audio_played_ms - user_speech_end_ms`, plus any transport leg you cannot see.

Be careful with vendor "e2e" numbers. LiveKit defines `e2e_latency` as time from the user stopping speaking to the agent beginning to respond. Retell's `latency.e2e` notes it excludes the network trip to the user's frontend. Neither includes the carrier leg on a PSTN call. Log your own milestones so you know what each number covers. Our latency guide breaks the budget down by stage.

Also log leading silence. Pipecat's `TTFAMetricsData` reports `ttfa`, `ttfb` and `leading_silence` separately. A fast TTFB with 300 ms of padded silence still feels slow.

Speech-to-text

FieldTypeWhy
`stt_request_id`stringDeepgram `metadata.request_id`; the first thing support will ask for
`stt_model_name` / `stt_model_version`stringDeepgram `model_info.name` and `model_info.version`
`transcript_final`string (redacted copy)What the agent acted on
`transcript_interims`arrayInterim revisions reveal instability and early endpointing
`is_final` / `speech_final` / `from_finalize`boolDeepgram segment flags; explain split utterances
`confidence`floatAlternative-level confidence
`words`array of word, start, end, confidenceWord timings anchor audio to text
`language` / `detected_languages`string, arrayCode-switching and wrong-language failures
`keyterms`array or hashDeepgram `keyterm` values in effect for this call
`endpointing_ms` / `utterance_end_ms`intSTT-side endpointing config
`audio_seconds`floatUsage for cost; Pipecat `STTUsageMetricsData`

Deepgram's docs warn not to treat `speech_final: true` alone as the full utterance. Long utterances can produce several `is_final: true` segments first. Log every final segment, not just the last. See Deepgram's endpointing and interim results guide.

Turn-taking

FieldTypeWhy
`turn_detection_mode`enum`vad`, `stt_endpointing`, `semantic_model`, `server_vad`
`turn_detection_params`mapSilence ms, thresholds, eagerness; hash it into `config_hash` too
`vad_events`arrayStart and stop pairs with timestamps
`interrupted`boolAgent speech was cut by the caller
`interruption_count`intPer turn and per call
`false_interruption`boolAgent stopped for a cough, "mm-hm" or background noise
`backchannel_count`intLiveKit `InterruptionMetrics.num_backchannels` when adaptive interruption is on
`overlap_ms`intBoth parties speaking at once
`dead_air_ms`intLongest silence with nobody speaking inside the turn
`agent_truncated_at_ms`intWhere playback was cut; OpenAI `conversation.item.truncate` carries `audio_end_ms`

Dead air and false interruptions are the two failures callers remember. Neither shows up in a transcript. Our endpointing guide covers the tuning side.

LLM and tool calls

FieldTypeWhy
`llm_response_id`stringOTel `gen_ai.response.id`; provider support needs it
`input_messages_ref`URI or hashFull context is large and sensitive; store by reference
`input_tokens` / `output_tokens` / `cached_input_tokens`intOTel `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.cache_read.input_tokens`
`finish_reason`string`stop`, `length`, `tool_calls`; `length` mid-sentence is a bug
`llm_ttft_ms` / `llm_duration_ms`intSplit thinking time from generation
`llm_retries` / `llm_fallback_used`int, boolFallbacks hide provider incidents
`tool_calls[].tool_call_id`stringJoin tool rows to turns
`tool_calls[].name` / `schema_version`stringFailure rate by tool and version
`tool_calls[].args_redacted`JSONArgument errors and hallucinated values
`tool_calls[].result_status`enum`ok`, `error`, `timeout`, `empty`, `invalid`
`tool_calls[].http_status` / `error_type`int, stringSeparate 4xx from 5xx from timeouts
`tool_calls[].latency_ms`intThe usual hidden latency spike
`tool_calls[].result_hash` / `result_ref`stringCompare what the tool said to what the agent claimed
`agent_claimed_outcome`string"Booked", "refunded"; checked against the tool result

Note the cached-token trap. The GenAI conventions say `gen_ai.usage.input_tokens` should already include cache reads. LiveKit's tracing docs make the same point. Adding the two double-counts.

The last row matters most. A tool can fail while the agent says "you're all set." That is a silent tool failure. You only catch it if you log the tool result next to the agent's claim.

Text-to-speech

FieldTypeWhy
`tts_request_id`stringProvider support
`tts_characters`intCost; LiveKit `characters_count`, Pipecat `metrics.character_count`
`tts_audio_seconds`floatLiveKit `audio_duration`
`tts_ttfb_ms` / `tts_leading_silence_ms`intPerceived latency
`tts_text_redacted`stringWhat the agent tried to say
`tts_error` / `tts_retries`string, intStalls and reconnects
`pronunciation_overrides`arrayLexicon entries used

Outcome

FieldTypeWhy
`end_reason`enumNormalized: `caller_hangup`, `agent_hangup`, `transfer`, `error`, `timeout`, `voicemail`
`end_reason_raw`stringVapi `endedReason`, Retell `disconnection_reason`, Twilio `CallStatus`
`who_hung_up`enum`caller`, `agent`, `system`
`transfer_target` / `transfer_result`string, enumRetell `transfer_destination`, `transfer_bridged` or `transfer_cancelled` events
`disposition`stringBusiness result code
`task_success_signal`enum`confirmed_by_system`, `claimed_only`, `failed`, `unknown`
`escalation_requested`boolCaller asked for a human
`compliance_flags`arrayDisclosures read, consent captured, opt-out honored
`last_speaker` / `trailing_silence_ms`enum, intWho spoke last and how long silence lasted before hangup; finds dead-air endings

Keep the raw vendor reason next to your normalized one. Vendor enums change, and Retell alone documents more than 30 disconnection reasons.

Cost per component

FieldUnitSource
`cost_telephony`USDTwilio `price` (connectivity only, posted after the call)
`cost_stt`USDaudio seconds x rate, or Deepgram request `usd`
`cost_llm`USDtokens x rate by model
`cost_tts`USDcharacters or audio seconds x rate
`cost_platform`USDVapi `costBreakdown` (`transport`, `stt`, `llm`, `tts`, `vapi`, `total`), Retell `call_cost.product_costs`
`cost_total`USDSum, with `pricing_version`

Store a `pricing_version` with each row. Rates change, and last quarter's cost per call must not silently recompute. This feeds cost per resolution.

Errors and warnings

Log every error as a structured event with `call_id`, `turn_id`, `component`, `error_type`, `provider_code`, `retryable` and `recovered`. Log warnings too: websocket reconnects, STT keepalive gaps, rate-limit headroom, fallback activations. OpenAI's Realtime API sends `rate_limits.updated` events, and its `error` event docs recommend logging error messages by default.

Quality scores added later

Evaluation writes back to the same keys. Keep scores in their own table so re-scoring never mutates the call record.

FieldTypeWhy
`call_id` / `turn_id`stringJoin key
`evaluator` / `evaluator_version`stringRubric or model version that produced the score
`metric`string`task_success`, `hallucination`, `interruption_handling`, `compliance`
`score` / `label` / `explanation`float, string, stringResult and rationale
`scored_at`timestampRe-scores are new rows, not updates

The GenAI conventions now include evaluation attributes such as `gen_ai.evaluation.name` and `gen_ai.evaluation.score.value`. They are in development status, so treat them as a naming guide rather than a stable contract.

Where each stack exposes these fields

You rarely need to invent these fields. Most stacks emit them. You need to collect them in one place with your IDs attached.

StackWhat it emitsHook or surface
LiveKit AgentsPer-turn `ChatMessage.metrics`: `transcription_delay`, `end_of_turn_delay`, `on_user_turn_completed_delay`, `llm_node_ttft`, `tts_node_ttfb`, `playback_latency`, `e2e_latency`, `started_speaking_at`, `stopped_speaking_at``conversation_item_added` event; `session_usage_updated`; `ctx.make_session_report()` in `on_session_end`; per-plugin `metrics_collected` with `speech_id`
LiveKit tracingOTel spans with `lk.` and `gen_ai.` attributes; content under `lk.pii.*` since Agents 1.7.0`set_tracer_provider(provider, metadata=..., allow_pii=False)`
Pipecat`MetricsFrame` with `TTFBMetricsData`, `TTFAMetricsData`, `TTFATMetricsData`, `ProcessingMetricsData`, `LLMUsageMetricsData`, `STTUsageMetricsData`, `TTSUsageMetricsData`, `TurnMetricsData``enable_metrics` and `enable_usage_metrics` in `PipelineParams`; observers such as `UserBotLatencyObserver`, `TurnTrackingObserver`, `ServiceMetricsObserver`
Pipecat tracingConversation, turn, stt, llm, tts spans; `turn.was_interrupted`, `metrics.ttfb`, `gen_ai.usage.*``setup_tracing(...)` plus `enable_tracing=True`, `enable_turn_tracking=True`, `conversation_id`
Deepgram streaming`Results` with `is_final`, `speech_final`, `from_finalize`, `confidence`, `words`, `metadata.request_id`, `metadata.model_info`; `Metadata`, `SpeechStarted`, `UtteranceEnd` messagesWebSocket messages; `tag` and `extra` params; Manage API `GET /v1/projects/{project_id}/requests`
OpenAI Realtime API`session.created`, `input_audio_buffer.speech_started` / `speech_stopped`, `conversation.item.input_audio_transcription.completed` / `failed`, `response.done` with `status`, `status_details`, `usage`, `rate_limits.updated`, `error`Server events over WebSocket or WebRTC
Twilio`CallSid`, `StreamSid`, `CallStatus`, `CallDuration`, `SipResponseCode`, `AnsweredBy`, `RecordingSid`; media `start`, `media`, `dtmf`, `mark`, `stop`Status callbacks; Media Streams WebSocket messages
Vapi`end-of-call-report` with `endedReason` and `artifact` (recording, transcript, messages); `status-update`, `speech-update`, `user-interrupted` with `turnId`, `model-output`, `hang`; call `costBreakdown` and `artifact.performanceMetrics`Server URL messages; call object
Retell`call_started`, `call_ended`, `call_analyzed`, `transcript_updated`, transfer events; call object with `agent_version`, `latency` (`e2e`, `asr`, `llm`, `tts`) percentiles, `call_cost`, `disconnection_reason`, `transcript_with_tool_calls`, `telephony_identifier.twilio_call_sid`Webhooks; `GET /v2/get-call/{call_id}`
OpenTelemetry GenAI`gen_ai.operation.name`, `gen_ai.provider.name`, `gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.response.id`, `gen_ai.response.finish_reasons`, `gen_ai.usage.*`, `gen_ai.tool.name`, `gen_ai.tool.call.id`Semantic conventions, now maintained in a separate GenAI repository

Sources: LiveKit data hooks, LiveKit OpenTelemetry traces, Pipecat metrics, Pipecat OpenTelemetry, Deepgram live audio reference, OpenAI Realtime conversations, Twilio Media Streams messages, Twilio Call resource, Vapi server events, Retell get-call, and the OpenTelemetry GenAI attribute registry. The GenAI conventions moved to the semantic-conventions-genai repository in 2026.

Two cautions. First, LiveKit notes that `llm_node_ttft` and `tts_node_ttfb` stay empty with a realtime model. Speech-to-speech stacks need different timing fields. Second, LiveKit Agents 1.7.0 renamed 12 span attributes to an `lk.pii.` prefix. Old dashboard queries return nothing and raise no error. Pin `framework_version` in your call envelope for exactly this reason. Running GPT-Live or another speech-to-speech model? Check its own event reference. The names above are verified for the Realtime API only.

Comparing frameworks? See LiveKit vs Pipecat latency and how to benchmark LiveKit vs Pipecat.

10 s
Retell webhook timeout before a retry, up to 3 retries
8,000 Hz
Twilio Media Streams audio: mu-law, one channel
5 min
Default LiveKit session_end_timeout for on_session_end work
7.5 s
Vapi response window for the assistant-request webhook

These are documented platform limits, and each one shapes your logging. Retell retries mean duplicate webhooks, so dedupe on `event` plus `call_id`. The 8 kHz narrowband audio caps what STT can do on phone calls. LiveKit's timeout bounds how much work fits in `on_session_end`. The Vapi window means you must never block on logging during call setup.

Example records: one call, one turn, one timeline

All values below are illustrative. Field names in our schema are recommendations. Provider-sourced fields keep the provider's names in comments.

Call envelope

{
  "call_id": "01J9ZK6T4Q8V3M2R7N5B1C0D9E",
  "started_at": "2026-09-24T15:02:11.482Z",
  "ended_at": "2026-09-24T15:06:47.903Z",
  "agent": {
    "agent_id": "billing-agent",
    "agent_version": "2026.09.24-3",
    "prompt_version": "billing/v41",
    "prompt_hash": "sha256:9f2c1e...",
    "config_hash": "sha256:41ab77...",
    "tool_schema_version": "tools/v12",
    "feature_flags": {"preemptive_generation": true},
    "framework": "livekit-agents",
    "framework_version": "1.7.2"
  },
  "models": {
    "stt": {"provider": "deepgram", "model": "nova-3", "version": "2026-08-12.1"},
    "llm": {"provider": "openai", "requested": "gpt-4.1", "served": "gpt-4.1-2025-04-14"},
    "tts": {"provider": "cartesia", "model": "sonic-2", "voice_id": "a0e99841"}
  },
  "telephony": {
    "direction": "inbound",
    "twilio_call_sid": "CA5d0d0d8047bf685c3f0ff980fe62c123",
    "twilio_stream_sid": "MZ18ad3ab5a668481ce02b83e7395059f0",
    "codec": "audio/x-mulaw",
    "sample_rate_hz": 8000,
    "sip_response_code": 200,
    "carrier": "carrier-a",
    "region": "us-east-1"
  },
  "outcome": {
    "end_reason": "caller_hangup",
    "end_reason_raw": "completed",
    "transfer_target": null,
    "disposition": "payment_plan_created",
    "task_success_signal": "confirmed_by_system"
  },
  "totals": {
    "turns": 14,
    "interruptions": 2,
    "false_interruptions": 1,
    "max_dead_air_ms": 1840,
    "p95_user_perceived_latency_ms": 1320
  },
  "cost_usd": {"telephony": 0.041, "stt": 0.021, "llm": 0.034, "tts": 0.052, "total": 0.148, "pricing_version": "2026-09"},
  "consent": {"recording": true, "basis": "announced_at_greeting", "jurisdiction": "US-CA"},
  "artifacts": {
    "audio_dual_channel": "s3://voice-raw/2026/09/24/01J9ZK.../call.wav",
    "transcript_redacted": "s3://voice-redacted/2026/09/24/01J9ZK.../transcript.json",
    "provider_report": "s3://voice-raw/2026/09/24/01J9ZK.../session_report.json"
  },
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "sampling": {"kept_reason": "always_keep:tool_error"}
}

Turn record

{
  "call_id": "01J9ZK6T4Q8V3M2R7N5B1C0D9E",
  "turn_id": "01J9ZK6T4Q8V3M2R7N5B1C0D9E:07",
  "turn_index": 7,
  "speech_id": "SP_7f3a...",
  "span_id": "00f067aa0ba902b7",
  "timing_ms": {
    "user_speech_start": 81230,
    "user_speech_end": 84910,
    "endpoint_decision": 85240,
    "stt_final": 85310,
    "llm_request": 85330,
    "llm_first_token": 85890,
    "first_tool_call": 86020,
    "tool_result": 87410,
    "tts_request": 87650,
    "tts_first_byte": 87790,
    "first_audio_played": 87860,
    "agent_speech_end": 91420
  },
  "derived_ms": {"user_perceived_latency": 2950, "tool_latency": 1390, "dead_air": 2950},
  "stt": {
    "request_id": "550e8400-e29b-41d4-a716-446655440000",
    "final_redacted": "yeah can I set up a payment plan for [AMOUNT]",
    "confidence": 0.93,
    "final_segments": 2,
    "interim_count": 6,
    "language": "en-US"
  },
  "llm": {
    "response_id": "resp_8a1f...",
    "input_tokens": 3120,
    "cached_input_tokens": 2816,
    "output_tokens": 64,
    "finish_reason": "tool_calls",
    "retries": 0
  },
  "tool_calls": [
    {
      "tool_call_id": "call_mszuSIzqtI65",
      "name": "create_payment_plan",
      "schema_version": "tools/v12",
      "args_redacted": {"account_id": "[ACCT]", "installments": 3},
      "result_status": "ok",
      "http_status": 200,
      "latency_ms": 1390
    }
  ],
  "tts": {"characters": 118, "ttfb_ms": 140, "leading_silence_ms": 0},
  "turn_taking": {"interrupted": false, "overlap_ms": 0, "false_interruption": false},
  "agent_claimed_outcome": "payment_plan_created"
}

This turn tells a story in one glance. Latency was 2.95 seconds, and 1.39 seconds of it was one tool. Without `tool_calls[].latency_ms`, you would blame the LLM.

Event timeline (compact)

[
  {"t_ms": 84910, "type": "vad.user_stop"},
  {"t_ms": 85240, "type": "turn.endpoint", "detector": "semantic", "delay_ms": 330},
  {"t_ms": 85310, "type": "stt.final", "request_id": "550e8400-...", "speech_final": true},
  {"t_ms": 85330, "type": "llm.request", "model": "gpt-4.1"},
  {"t_ms": 85890, "type": "llm.first_token"},
  {"t_ms": 86020, "type": "tool.start", "tool_call_id": "call_mszuSIzqtI65", "name": "create_payment_plan"},
  {"t_ms": 87410, "type": "tool.end", "status": "ok", "http_status": 200},
  {"t_ms": 87650, "type": "tts.request", "chars": 118},
  {"t_ms": 87790, "type": "tts.first_byte"},
  {"t_ms": 87860, "type": "playback.start", "twilio_mark": "t07-seg1"},
  {"t_ms": 91420, "type": "playback.end", "twilio_mark": "t07-seg1"}
]

`t_ms` is an offset from call start on one monotonic clock. That removes cross-service clock skew from every duration you compute.

Implementation: structured logs and OpenTelemetry spans per turn

The pattern is simple. Build one in-memory turn object. Stamp milestones on a monotonic clock. Emit one structured log line and one span when the turn closes. Never do network I/O on the audio path.

The code below is simplified for readability. It uses the standard OpenTelemetry Python API and the standard library's `QueueHandler`. Adapt names to your schema.

import json, logging, logging.handlers, queue, time, uuid
from dataclasses import dataclass, field, asdict
from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode

# 1) Non-blocking logging: the audio loop only enqueues.
log_queue: queue.Queue = queue.Queue(maxsize=50_000)
root = logging.getLogger("voice")
root.setLevel(logging.INFO)
root.addHandler(logging.handlers.QueueHandler(log_queue))

class JsonFormatter(logging.Formatter):
    def format(self, record):
        return json.dumps(getattr(record, "payload", {"msg": record.getMessage()}))

sink = logging.StreamHandler()          # swap for your shipper (file, OTLP logs, Kafka)
sink.setFormatter(JsonFormatter())
listener = logging.handlers.QueueListener(log_queue, sink)
listener.start()                         # I/O happens on the listener thread

tracer = trace.get_tracer("voice-agent")

@dataclass
class Turn:
    call_id: str
    index: int
    t0_ns: int                           # call start, monotonic
    marks: dict = field(default_factory=dict)
    tool_calls: list = field(default_factory=list)
    attrs: dict = field(default_factory=dict)

    def mark(self, name: str) -> None:
        self.marks[name] = (time.monotonic_ns() - self.t0_ns) // 1_000_000

    @property
    def turn_id(self) -> str:
        return f"{self.call_id}:{self.index:02d}"

class CallLogger:
    def __init__(self, envelope: dict):
        self.call_id = envelope.setdefault("call_id", str(uuid.uuid4()))
        self.t0_ns = time.monotonic_ns()
        self.envelope = envelope
        self.envelope["started_at_unix_ms"] = time.time_ns() // 1_000_000  # one wall-clock anchor

    def new_turn(self, index: int) -> Turn:
        return Turn(self.call_id, index, self.t0_ns)

    def close_turn(self, turn: Turn) -> None:
        m = turn.marks
        upl = m.get("first_audio_played", 0) - m.get("user_speech_end", 0) if "first_audio_played" in m else None
        with tracer.start_as_current_span("voice.turn") as span:
            span.set_attribute("voice.call_id", turn.call_id)
            span.set_attribute("voice.turn_id", turn.turn_id)
            span.set_attribute("voice.agent_version", self.envelope["agent"]["agent_version"])
            span.set_attribute("voice.prompt_hash", self.envelope["agent"]["prompt_hash"])
            if upl is not None:
                span.set_attribute("voice.user_perceived_latency_ms", upl)
            for k in ("gen_ai.request.model", "gen_ai.response.model",
                      "gen_ai.usage.input_tokens", "gen_ai.usage.output_tokens"):
                if k in turn.attrs:
                    span.set_attribute(k, turn.attrs[k])
            failed = [t for t in turn.tool_calls if t["result_status"] != "ok"]
            if failed:
                span.set_status(Status(StatusCode.ERROR, "tool_failure"))
            ctx = span.get_span_context()
            payload = {
                "kind": "turn", "call_id": turn.call_id, "turn_id": turn.turn_id,
                "trace_id": format(ctx.trace_id, "032x"), "span_id": format(ctx.span_id, "016x"),
                "timing_ms": m, "user_perceived_latency_ms": upl,
                "tool_calls": turn.tool_calls, **turn.attrs,
            }
        root.info("turn", extra={"payload": payload})   # enqueue only

async def logged_tool(turn: Turn, name: str, schema_version: str, fn, args_redacted: dict, **kwargs):
    """Wrap every tool so failures and latency are never invisible."""
    start = time.monotonic_ns()
    if "first_tool_call" not in turn.marks:
        turn.mark("first_tool_call")
    rec = {"name": name, "schema_version": schema_version, "args_redacted": args_redacted}
    with tracer.start_as_current_span("execute_tool " + name) as span:
        span.set_attribute("gen_ai.operation.name", "execute_tool")
        span.set_attribute("gen_ai.tool.name", name)
        try:
            result = await fn(**kwargs)
            rec["result_status"] = "empty" if result in (None, {}, []) else "ok"
            return result
        except TimeoutError as e:
            rec.update(result_status="timeout", error_type=type(e).__name__)
            span.record_exception(e); span.set_status(Status(StatusCode.ERROR))
            raise
        except Exception as e:
            rec.update(result_status="error", error_type=type(e).__name__)
            span.record_exception(e); span.set_status(Status(StatusCode.ERROR))
            raise
        finally:
            rec["latency_ms"] = (time.monotonic_ns() - start) // 1_000_000
            turn.tool_calls.append(rec)
            turn.mark("tool_result")

Three details carry most of the value. The tool wrapper classifies `empty` separately from `ok`, which is how silent failures surface. Durations come from one monotonic clock per call. Only the listener thread touches the network.

Hooking the timings in LiveKit Agents

LiveKit already computes per-turn latency. Read it from the chat item and attach your IDs. The event and field names below come from the LiveKit data hooks docs. The mapping into `CallLogger` is ours and simplified.

from livekit.agents import ConversationItemAddedEvent, SessionUsageUpdatedEvent
from livekit.agents.llm import ChatMessage

def wire_livekit(session, call_log: CallLogger):
    @session.on("conversation_item_added")
    def _on_item(ev: ConversationItemAddedEvent):
        if not isinstance(ev.item, ChatMessage):
            return
        m = ev.item.metrics or {}
        root.info("lk_turn_metrics", extra={"payload": {
            "kind": "turn_metrics", "call_id": call_log.call_id, "role": ev.item.role,
            "transcription_delay": m.get("transcription_delay"),
            "end_of_turn_delay": m.get("end_of_turn_delay"),
            "llm_node_ttft": m.get("llm_node_ttft"),
            "tts_node_ttfb": m.get("tts_node_ttfb"),
            "e2e_latency": m.get("e2e_latency"),
            "started_speaking_at": m.get("started_speaking_at"),
            "stopped_speaking_at": m.get("stopped_speaking_at"),
        }})

    @session.on("session_usage_updated")
    def _on_usage(ev: SessionUsageUpdatedEvent):
        for u in ev.usage.model_usage:
            root.info("lk_usage", extra={"payload": {
                "kind": "usage", "call_id": call_log.call_id,
                "provider": u.provider, "model": u.model, "usage": str(u)}})

# At session end, persist the SDK's own report as an artifact:
async def on_session_end(ctx):
    report = ctx.make_session_report().to_dict()
    # upload report to object storage off the hot path; store its URI on the call envelope

For spans, call `set_tracer_provider(provider, metadata={"voice.call_id": call_id}, allow_pii=False)` before the session starts. The `metadata` lands on every span. `allow_pii=False` strips `lk.pii.*` and GenAI content attributes from your exporters. Deeper wiring lives in our OpenTelemetry guide for voice agents.

Hooking the timings in Pipecat

Pipecat exposes metrics as frames and observers. The observer names, event names and handler signatures below come from the Pipecat metrics docs. As of September 2026, the docs use `PipelineWorker`; older releases used `PipelineTask`.

from pipecat.observers.user_bot_latency_observer import UserBotLatencyObserver
from pipecat.observers.turn_tracking_observer import TurnTrackingObserver
from pipecat.observers.service_metrics_observer import ServiceMetricsObserver
from pipecat.pipeline.worker import PipelineWorker, PipelineParams

latency_obs = UserBotLatencyObserver()
turn_obs = TurnTrackingObserver(turn_end_timeout_secs=2.5)
svc_obs = ServiceMetricsObserver()

@latency_obs.event_handler("on_latency_measured")
async def _lat(observer, latency_seconds):
    root.info("pc_latency", extra={"payload": {"kind": "user_bot_latency",
              "call_id": CALL_ID, "ms": int(latency_seconds * 1000)}})

@turn_obs.event_handler("on_turn_ended")
async def _turn(observer, turn_count, duration, was_interrupted):
    root.info("pc_turn", extra={"payload": {"kind": "turn_end", "call_id": CALL_ID,
              "turn_index": turn_count, "duration_s": duration, "interrupted": was_interrupted}})

@svc_obs.event_handler("on_service_latency")
async def _svc(observer, record):
    root.info("pc_svc", extra={"payload": {"kind": "service_latency", "call_id": CALL_ID,
              "metric": record.kind, "seconds": record.seconds, "processor": record.processor}})

worker = PipelineWorker(
    pipeline,
    params=PipelineParams(enable_metrics=True, enable_usage_metrics=True),
    observers=[latency_obs, turn_obs, svc_obs],
    enable_tracing=True,
    enable_turn_tracking=True,
    conversation_id=CALL_ID,                          # your call_id becomes conversation.id
    additional_span_attributes={"voice.agent_version": AGENT_VERSION},
)

Passing your call_id as `conversation_id` is the single most useful line. Every Pipecat span then joins to your warehouse rows without a lookup table.

Joining provider IDs

Capture provider IDs at the moment they appear. Do not reconstruct them later.

1. On the Twilio Media Streams `start` message, read `start.callSid`, `start.streamSid` and `start.customParameters`. Put your call_id in the custom parameters when you build the ``.

2. On every Deepgram `Results` message, log `metadata.request_id` with the turn. Set `tag` to your agent ID for usage reporting.

3. On Retell, read `telephony_identifier.twilio_call_sid` from the call object. On Vapi, the deprecated `phoneCallProviderId` still carries the Twilio callSid for phone calls.

4. On each LLM response, log the response ID. It maps to `gen_ai.response.id`.

5. On webhooks, dedupe. Retell retries up to 3 times after a 10-second timeout. Key per-call events on `event` plus `call_id`.

Logging must never add latency

Every millisecond spent logging on the audio path is a millisecond of dead air. Follow four rules.

  • Enqueue, never send. The hot path writes to an in-memory queue. A separate thread or process ships it.
  • Batch spans. Use the SDK's `BatchSpanProcessor`, as both LiveKit and Pipecat examples do.
  • Drop debug, never errors. If the queue fills, shed interim and debug events first. Keep turn, tool and error records.
  • Upload artifacts after the call. Audio and reports go to object storage once the session ends. Respect LiveKit's `session_end_timeout` or raise it.

Monitoring tool call failures and latency spikes from these logs

This is the most common question we get from teams running agents in production. The answer is three alerts built on the turn and tool tables.

First, alert on tool failure rate by tool and schema version. Count `timeout`, `error`, `empty` and `invalid` together. Silent failures hide inside `empty`.

Second, alert on p95 user-perceived latency, not mean. Break it down by stage so the alert names a suspect: endpointing, STT, LLM, tool, or TTS.

Third, alert on claim-versus-result mismatches. When `agent_claimed_outcome` says "booked" and the tool status is not `ok`, page someone.

For the broader alerting design, see how to monitor AI voice agents in production. For alert payloads, see our walkthrough of per-turn call monitoring alerts.

Deepgram usage and performance: what to log and where to pull it

Log three things on every Deepgram stream: `metadata.request_id`, `metadata.model_info` (name and version), and audio seconds sent. Add STT timing from your turn record. That gives you performance.

For usage, Deepgram's Manage API lists requests per project at `GET /v1/projects/{project_id}/requests`. You can filter by `request_id`, `endpoint`, `method` (`streaming`) and `status` (`succeeded` or `failed`). Each request carries a response code and usage details. Reconcile that daily against your own audio-seconds totals. A gap means dropped streams or billing surprises. See the Deepgram requests reference.

Versioning prompts, test cases, assertions and results

Treat prompts like code and results like data. Store the prompt text in version control. Compute `prompt_hash` from the exact rendered template in CI. Stamp both on every call.

Give test cases stable IDs and a version. An assertion is a small typed rule tied to a test case. One example: "tool `create_payment_plan` called once with installments 3." Results are rows keyed by `test_case_id`, `test_case_version`, `agent_version`, `prompt_hash` and `run_id`.

With that shape, prompt regressions become a query. Compare pass rates and production metrics before and after a `prompt_hash` change. Our post on voice agent metric drift covers the statistics.

Can you reconstruct any call in five minutes?
Evalgent turns your call logs into independent, repeatable evaluation of every call.
Book a demo

Storage, sampling and retention

Different layers have different shapes. Store them accordingly.

A voice agent logging pipeline: events emitted asynchronously from the agent, PII redaction before storage, hot event storage and cold audio storage, then dashboards, alerts, replay and evaluation
TierDataStoreRetention (example)Access
HotCall and turn rows, errors, recent eventsWarehouse or columnar log store30 to 90 days queryableEngineering, on-call
WarmFull event timelines, spansTrace backend or partitioned tables14 to 30 daysEngineering
ColdRedacted transcripts, provider reportsObject storage, lifecycle rulesPer policy, often 1 yearEval and QA
RestrictedRaw audio, raw transcriptsSeparate encrypted bucketShortest your policy allowsNamed roles, audited

Retention periods are examples, not recommendations for your jurisdiction. Your legal and compliance owners set them.

Sampling rules

Sample volume, never failures. Keep 100 percent of these calls:

  • Any error, tool failure, fallback or retry
  • Any escalation, transfer or "talk to a human" request
  • Any compliance-relevant call: payments, health, consent, opt-out
  • Any call with latency or dead air above threshold
  • Any call touched by an experiment or new version in its first days

Sample the rest for full event timelines if cost demands it. Keep call and turn rows for every call. They are small.

Index for the queries you will run

Partition by date. Cluster or index by `agent_version`, `prompt_hash`, `carrier`, `region` and `end_reason`. Put `call_id` on everything. Store tool calls in their own table, keyed by `turn_id`. Then failure rates need no JSON unnesting at query time.

Privacy and compliance: redact before you persist

Logging makes calls debuggable. It also creates a copy of everything the caller said. Design for that from day one. This section is general engineering guidance, not legal advice.

The core rule: redact before persistence, and keep raw and redacted data physically separate. Your default pipeline writes only redacted transcripts and hashed identifiers. Raw audio and raw text go to a restricted store, if they are kept at all.

A practical pipeline looks like this:

1. Transcribe in memory. Run entity detection on the final transcript before it is written anywhere.

2. Replace entities with typed placeholders such as `[CARD]`, `[DOB]`, `[ACCT]`. Keep a reversible mapping only if a named process needs it.

3. Strip content from spans. LiveKit's `allow_pii=False` and the `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=false` setting keep prompts and tool payloads out of trace exporters.

4. Suppress DTMF digits in payment flows. Log that digits were pressed, never which.

5. Mute or pause recording during card capture, or route payments to a compliant IVR segment.

6. Attach consent flags to the call envelope: whether recording was announced, how consent was captured, and the caller's likely jurisdiction.

7. Support deletion by `call_id` and by hashed caller identifier across every tier, including object storage and backups.

On card data, the PCI Security Standards Council is clear. Sensitive authentication data, including card verification codes, must not be stored after authorization. Its FAQ on SAD storage explains why this holds even without a PAN present. Audio and transcripts count, so a caller reading a CVV aloud is a logging problem.

On health data, see the HIPAA Privacy Rule's minimum necessary standard. It asks covered entities to limit PHI use and access to what a purpose requires. For logs, that means role-based access and redacted defaults.

On recording consent, US law varies by state. The Digital Media Law Project's guide to recording phone calls summarizes one-party and all-party consent regimes. Its page notes it is no longer updated, so confirm current rules with counsel. Log what consent you obtained so you can prove it later.

Platform settings help but do not replace your pipeline. Retell's `data_storage_setting` supports `everything_except_pii` and `basic_attributes_only`. When storage is opted out, Retell's webhook docs say the `recording_url` expires after 10 minutes. Download it inside that window or lose it. For the full redaction playbook, see our PII handling guide for voice agents.

The 10 queries your logs must answer

If your schema cannot answer these, fix the schema. SQL below is generic and simplified. Table names match the schema above.

1. p95 user-perceived latency by carrier, last 7 days

SELECT c.carrier,
       APPROX_QUANTILES(t.user_perceived_latency_ms, 100)[OFFSET(95)] AS p95_ms,
       COUNT(*) AS turns
FROM turns t JOIN calls c USING (call_id)
WHERE c.started_at >= CURRENT_DATE - 7
GROUP BY c.carrier
ORDER BY p95_ms DESC;

2. Tool failure rate by tool and schema version

SELECT name, schema_version,
       COUNTIF(result_status IN ('error','timeout','empty','invalid')) / COUNT(*) AS failure_rate,
       COUNT(*) AS calls
FROM tool_calls
WHERE ts >= CURRENT_DATE - 1
GROUP BY name, schema_version
ORDER BY failure_rate DESC;

3. Interruption rate before and after a prompt change

SELECT c.prompt_hash,
       AVG(CAST(t.interrupted AS INT64)) AS interruption_rate,
       AVG(CAST(t.false_interruption AS INT64)) AS false_interruption_rate
FROM turns t JOIN calls c USING (call_id)
WHERE c.agent_id = 'billing-agent' AND c.started_at >= '2026-09-17'
GROUP BY c.prompt_hash;

4. Calls that ended in dead air

SELECT call_id, max_dead_air_ms, end_reason
FROM calls
WHERE end_reason = 'caller_hangup'
  AND last_speaker = 'caller'
  AND trailing_silence_ms > 5000;

5. Cost per resolved call by agent version

SELECT agent_version,
       SUM(cost_total) / NULLIF(COUNTIF(task_success_signal = 'confirmed_by_system'), 0) AS cost_per_resolution
FROM calls
WHERE started_at >= CURRENT_DATE - 30
GROUP BY agent_version;

6. Regressions after an LLM snapshot change

SELECT llm_model_served,
       AVG(CAST(task_success_signal = 'confirmed_by_system' AS INT64)) AS success_rate,
       APPROX_QUANTILES(p95_user_perceived_latency_ms, 100)[OFFSET(50)] AS median_call_p95_ms
FROM calls
WHERE agent_id = 'billing-agent' AND started_at >= CURRENT_DATE - 14
GROUP BY llm_model_served;

7. Claim-versus-result mismatches (silent tool failures)

SELECT t.call_id, t.turn_id, t.agent_claimed_outcome, tc.name, tc.result_status
FROM turns t JOIN tool_calls tc USING (turn_id)
WHERE t.agent_claimed_outcome IS NOT NULL
  AND tc.result_status <> 'ok';

8. Which stage dominates slow turns (uses derived per-turn duration columns)

SELECT CASE
         WHEN tool_latency_ms > llm_ttft_ms AND tool_latency_ms > tts_ttfb_ms THEN 'tool'
         WHEN llm_ttft_ms > tts_ttfb_ms THEN 'llm'
         ELSE 'tts' END AS dominant_stage,
       COUNT(*) AS slow_turns
FROM turns
WHERE user_perceived_latency_ms > 2000
GROUP BY dominant_stage;

9. STT errors and reconnects by Deepgram model version

SELECT stt_model_version, COUNT(DISTINCT call_id) AS affected_calls
FROM events
WHERE type IN ('stt.error', 'stt.reconnect')
GROUP BY stt_model_version;

10. Eval scores joined back to versions

SELECT c.agent_version, e.metric, AVG(e.score) AS avg_score, COUNT(*) AS n
FROM eval_results e JOIN calls c USING (call_id)
WHERE e.evaluator_version = 'rubric-v7'
GROUP BY c.agent_version, e.metric;

`APPROX_QUANTILES` and `COUNTIF` are BigQuery syntax. In Postgres, use `percentile_cont(0.95) WITHIN GROUP (ORDER BY ...)` and `COUNT(*) FILTER (WHERE ...)`.

Anti-patterns that make calls impossible to debug

We see the same six mistakes across stacks.

Logging only transcripts. A transcript cannot show latency, overlap, dead air or tone. It also shows what STT heard, not what the caller said.

No version info. Without `prompt_hash`, `llm_model_served` and `config_hash`, you cannot attribute a regression. You can only notice it.

Wall-clock skew across services. Subtracting timestamps from three hosts produces negative latencies and phantom spikes. Use one monotonic clock per call and store offsets.

Logging audio without consent. Recording first and asking later creates legal exposure. It also poisons your eval dataset if you must delete it.

Sampling away the failures. Uniform 10 percent sampling keeps 10 percent of your incidents. Always keep errors, escalations and compliance calls.

Unstructured strings. `print(f"tool failed: {e}")` cannot be counted, grouped or alerted on. Emit typed fields.

More failure patterns live in why voice agents fail in production.

How logs feed evaluation

Good logs turn every production call into test material. Three loops matter.

Replay as regression tests. Take calls that failed, strip them to caller audio plus expected outcomes, and replay them against each new version. Dual-channel audio and turn timing make replay realistic. Pair this with synthetic callers for coverage and shadow testing for live traffic.

Score every call, not a sample. Rubric and model-based evaluators can score each call on task success, compliance and conversational quality. Scores write to `eval_results` keyed by `call_id` and `turn_id`.

Link scores back to versions. Because every call carries its versions, a score drop points at a specific change. Curate the best-labeled calls into a golden dataset.

This is where independence matters. The team that built the agent should not be the only one grading it. Evalgent is an independent evaluation platform for voice agents. It ingests the call, turn, event and artifact data described here. It scores every call against your rubrics and writes results back by `call_id`. The schema in this post is the shape we ask teams to send. For how this fits the wider stack, see voice agent observability.

How to roll out call logging in a week

A one-week plan for a team with an agent already in production.

1. Day 1: Define IDs and the envelope. Generate `call_id` at the first touchpoint. Add `agent_version`, `prompt_hash`, `config_hash` and model fields. Ship it to one table.

2. Day 2: Capture provider IDs. Log Twilio `CallSid` and `StreamSid`, Deepgram `request_id`, LLM response IDs, and your platform's call ID. Push your call_id into provider metadata.

3. Day 3: Add per-turn timing. Wire LiveKit `conversation_item_added` or Pipecat observers. Stamp your own milestones on a monotonic clock. Compute user-perceived latency.

4. Day 4: Wrap every tool. Add the tool wrapper with status, latency, error type and redacted args. Log `agent_claimed_outcome` next to tool results.

5. Day 5: Redact and split storage. Add entity redaction before writes. Create separate raw and redacted buckets with different access roles. Add consent flags.

6. Day 6: Build the queries and alerts. Implement the 10 queries. Alert on tool failure rate, p95 latency and claim mismatches.

7. Day 7: Run the five-minute test. Pick five random calls. Time each reconstruction. Fix whatever blocked you, then connect scoring.

Copyable checklist

Paste this into your logging design review. Every line should be true before you call logging done.

  • `call_id` generated at first touchpoint and present on every record
  • `turn_id` on every turn, event, tool call and eval score
  • `agent_version`, `prompt_version`, `prompt_hash`, `config_hash`, `tool_schema_version` on the envelope
  • STT, LLM and TTS provider, model and served version logged
  • Twilio `CallSid` and `StreamSid`, Deepgram `request_id`, LLM response IDs captured
  • Per-turn milestones on one monotonic clock, with user-perceived latency derived
  • Endpointing mode and parameters, interruptions, false interruptions, overlap and dead air logged
  • Every tool call logged with status, latency, error type and redacted args
  • Agent-claimed outcome logged next to the tool result
  • Cost per component with a `pricing_version`
  • Raw end reason and normalized end reason both stored
  • Logging is queue-based and never blocks the audio path
  • Redaction runs before persistence; raw and redacted stores are separate
  • DTMF digits suppressed in payment flows; consent flags on the envelope
  • 100 percent retention for errors, escalations and compliance calls
  • Deletion by `call_id` works across every tier
  • Eval results stored in their own table, keyed by `call_id` and `turn_id`
  • Five-minute reconstruction test passes on random calls

The bottom line

Voice agent call logging is the foundation that testing, evaluation and debugging stand on. Teams that log four joined layers, with versions and provider IDs on every record and redaction before storage, can reconstruct, replay and score any call.

Frequently asked questions

How do I monitor tool call failures and latency spikes in production voice agents?

Monitor tool call failures and latency spikes by wrapping every tool call. Log status, latency, error type and schema version, keyed to a turn_id. Alert on failure rate by tool and version, counting timeouts and empty results. Alert on p95 user-perceived latency broken down by stage. Also alert when the agent claims success but the tool result failed.

How do I monitor voice agents in production?

To monitor voice agents in production, log a call envelope, per-turn timing, tool results and outcomes for every call. Build dashboards on p95 user-perceived latency, tool failure rate, interruption rate, dead air and task success. Alert on thresholds per agent version. Then score calls with an independent evaluator so quality regressions surface, not just outages.

How do I monitor and log Deepgram API usage and performance in production?

Log Deepgram's metadata.request_id, model_info name and version, and audio seconds on every stream. Record STT timing per turn, including when speech_final arrives. For usage, pull the Manage API's project requests endpoint, filtered by streaming method and status. Reconcile it daily against your own audio totals to catch dropped streams or billing gaps.

How do I detect prompt regressions in a voice agent?

Detect prompt regressions by stamping a prompt_hash and prompt_version on every call. Compare task success, interruption rate, tool failure rate and latency before and after each hash change. Replay a fixed regression suite of past failed calls against every new prompt. Without the hash on each call, you can notice a regression but never attribute it.

How do I version voice agent prompts, test cases, assertions, and results?

Version voice agent prompts in source control and stamp a hash of the rendered prompt on every call. Give test cases stable IDs and versions. Write assertions as typed rules tied to a test case. Store results as rows keyed by test case version, agent version, prompt hash and run ID. Then any result traces to exact inputs.

Should I record audio on every voice agent call?

Record audio on every call only where consent and policy allow. Dual-channel recordings make replay, overlap and barge-in analysis possible, which transcripts cannot. Announce recording, log how consent was captured, and pause recording during card capture. Store raw audio in a restricted, encrypted bucket with short retention. This is general guidance, not legal advice.

What is user-perceived latency in a voice agent?

User-perceived latency is the time from the end of the caller's speech to the first agent audio reaching them. It spans endpointing, STT, LLM time to first token, tool calls and TTS first byte. Vendor "e2e" numbers often exclude network or carrier legs, so log your own milestones and derive it yourself.

How long should I retain voice agent call logs?

Voice agent call log retention depends on your legal and compliance requirements. A common pattern keeps call and turn rows queryable for 30 to 90 days. Redacted transcripts stay longer in cold storage. Raw audio gets the shortest retention your policy allows. Keep deletion by call_id working everywhere.

Want every call scored independently, straight from these logs? Book a demo.

Related Articles