Evalgent
Back to Blog
Voice AI Evaluation

Voice agent testing vs monitoring vs observability (and evaluation): a working model

Deepesh Jayal
Updated
24 min read
Voice agent testing vs monitoring vs observability (and evaluation): a working model
On this page

Most articles on this topic stop at a four-row table. That table is correct and nearly useless. It tells you the words mean different things. It does not tell you which practice catches an endpointing bug, how many calls you must score to see a one-point regression, or how to page someone when task success drops.

This guide does those things. It defines each practice by the question it answers and the data it needs. It maps seven voice failure layers to the practice that catches each one, and when. It shows a voice turn as an OpenTelemetry trace using the current GenAI conventions. Then it works through the sampling math, the alert design, a maturity model, and a runbook you can copy.

pass^8 < 25%
gpt-4o consistency on tau-bench retail tasks (Yao et al., 2024)
~1%
Whisper transcripts with fully hallucinated phrases (Koenecke et al., 2024)
14.4x
1-hour burn rate that pages at 2% budget spend (Google SRE Workbook)

The four practices, defined by the question each answers

The cleanest split is not by tool or by team. It is by the question you are asking and the data you need to answer it.

PracticeThe question it answersKind of unknownData it needsWhen it runsOutput
TestingDoes this build behave as specified under conditions we chose?Known knownsScripted scenarios, synthetic caller audio, expected outcomesBefore deploy, on every changePass or fail per scenario
EvaluationHow good was this call, against a written rubric?Defined criteriaAudio, transcript, tool calls, rubricBoth sides of launchScores and labels
MonitoringIs a known indicator outside its bound right now?Known unknownsAggregated time series of scores and system metricsAfter deploy, continuouslyAlerts and tickets
ObservabilityWhy did this specific call go the way it did?Unknown unknownsRaw, wide, high-cardinality events and tracesAfter deploy, on demandExplanations and new test cases

The "known unknowns" language comes from Charity Majors. In her observability manifesto she puts it plainly: "Monitoring is about known-unknowns and actionable alerts." Observability is the ability to ask new questions of your system without shipping new code to collect new data. Her test is concrete. If you pre-aggregate, or you cannot group by high-cardinality fields like a request ID, you are not doing observability.

Google's SRE book chapter on monitoring adds the other half. A monitoring system should answer two questions: what is broken, and why. The "what" is the symptom. The "why" is a cause. The book recommends paging on symptoms and keeping cause-hunting for debugging. That maps cleanly: monitoring pages on symptoms, observability finds causes.

Evaluation is the odd one out because it is not a phase. It is a function. You feed it a call and a rubric, and it returns a score. Testing calls that function on synthetic calls before launch. Monitoring calls it on real calls after launch and watches the scores as a time series. Observability uses the scores as one more field to slice by. If you build one evaluator and reuse it everywhere, the four practices become one system.

The voice-specific problem: the success response that lies

The SRE book lists three kinds of errors. Some are explicit, like an HTTP 500. Some are by policy, like a response slower than your commitment. The third kind is implicit: a success response carrying the wrong content. The book notes that only end-to-end system tests can detect that you are serving the wrong content.

That third kind is the dominant failure mode of a voice agent. The caller asks to move an appointment to Tuesday. The agent says "Done, you're all set for Thursday." Every span is green. STT returned a final transcript. The LLM returned in 600 ms. The tool returned 200. TTS played clean audio. No infrastructure metric moved.

This is why voice teams cannot copy a web-service monitoring playbook. The golden signals of latency, traffic, errors and saturation still matter. But the error that hurts most is semantic, and only an evaluator can turn it into a number. Once it is a number, monitoring can watch it and alert on it. Evaluation is the bridge that lets the other three practices see meaning.

Why pre-deploy testing cannot prove a low failure rate

Teams often treat a green test suite as proof that the agent is reliable. The statistics say otherwise, and this is the main reason testing can never replace monitoring.

The rule of three

Suppose you run 300 synthetic calls on a release candidate and all 300 pass. What failure rate have you ruled out? Hanley and Lippman-Hand answered this in their 1983 JAMA paper If Nothing Goes Wrong, Is Everything All Right?. When you observe zero events in n trials, the 95% upper bound on the true rate is about 3/n.

Clean test runs (zero failures)95% upper bound on true failure rate
1003.0%
3001.0%
1,0000.3%
3,0000.1%

So 300 clean runs only tell you the failure rate is probably under 1%. A compliance failure at 0.3% of calls is invisible to that suite. At 10,000 calls a day, 0.3% is 30 bad calls a day. You need production volume to see failures that rare.

Single-run pass rates overstate reliability

The tau-bench paper from Yao and colleagues made a related point for tool-using agents. They simulated users with an LLM, gave the agent real API tools and policy rules, and checked the database state at the end of each conversation. They also introduced pass^k: the chance an agent succeeds on the same task in all k independent trials.

Their headline: gpt-4o solved under 50% of tasks, and its pass^8 on the retail domain fell below 25%. An agent that passes a scenario once may fail the same scenario on the next run. For a voice agent, where ASR noise and sampling add variance on top of the LLM, this matters even more.

If trials were independent, pass^k is just p to the power k. This illustrative table shows how fast consistency decays:

Per-run success (p)pass^1pass^4pass^8
0.9090%66%43%
0.9595%81%66%
0.9999%96%92%

The practical rule: run each critical scenario several times and gate on the worst run, not the average. Our guide to synthetic callers for voice agent testing covers how to vary the caller between runs.

What testing is actually for

None of this makes testing less important. It makes its job clearer. Testing catches deterministic breaks and large regressions before any caller hears them. A prompt edit that breaks date parsing fails every time. A model swap that doubles the transfer rate shows up in 200 runs. What testing cannot do is measure small rates. That job belongs to evaluation running on production traffic, watched by monitoring.

A failure taxonomy: which practice catches what, and when

A voice agent fails at seven layers. Each layer fails in its own way, and each is caught first by a different practice. This matrix is the core of the guide.

LayerTypical failureCaught pre-deploy byCaught post-deploy byExplained by these fields
ASR / STTWrong entity, hallucinated words on silence, accent gapsTesting with noisy and accented audio, scored on entity accuracyMonitoring ASR confidence, re-ask rate, per-segment entity failureAudio clip, transcript, confidence, STT model and version
Endpointing / turn-takingAgent cuts in mid-sentence, or waits too longTesting with scripted pauses and barge-insMonitoring interruption rate and turn latencyVAD start and stop times, end-of-turn event, silence timer setting
LLM reasoningWrong answer delivered fluently, lost contextEvaluation of test calls against a rubricEvaluation of production calls, watched as a ratePrompt version, model ID, full context, tool results
Tool / integrationWrong arguments, empty payload returned with 200, timeoutTesting against a stubbed backend with assertionsMonitoring tool error and empty-result ratesTool name, arguments, result, duration
TTSMispronounced names, numbers read wrong, clipped audioTesting with entity-heavy prompts, listened or scoredMonitoring TTS time to first byte and character countsText sent to TTS, voice ID, audio clip
Telephony / transportOne-way audio, codec mismatch, dropped callsLimited: a few carrier test callsMonitoring call setup failure, silence-only calls, jitterCarrier, codec, SIP response codes, region
Policy / complianceMissing disclosure, wrong refund promise, PII read aloudEvaluation of adversarial test callsEvaluation of every production callTranscript, policy rule ID, agent version
Matrix mapping seven voice agent failure layers to testing, evaluation, monitoring and observability, showing which practice catches each failure first and whether before or after deployment

What the matrix teaches

Three patterns stand out.

First, telephony failures are almost invisible before launch. You can place a few test calls, but you cannot reproduce every carrier, region and handset. PSTN audio arrives as 8 kHz narrowband, often G.711, and your pipeline may resample it to 16 kHz for STT. A codec mismatch or a one-way audio bug shows up as a call with caller speech but no agent audio, or the reverse. Only monitoring at production scale catches the long tail here.

Second, ASR failures need audio in the trace. The Careless Whisper study by Koenecke and colleagues found that roughly 1% of Whisper transcriptions contained entire hallucinated phrases with no basis in the audio. 38% of those hallucinations contained explicit harms. Hallucinations were more frequent for speakers with longer non-vocal stretches. In a voice agent, that means long caller pauses are where invented words appear. A transcript-only log cannot show this. You need the audio clip next to the transcript to see that the words were never spoken.

Third, benchmark accuracy does not transfer. In WER We Are and WER We Think We Are, Szymański and colleagues compared commercial ASR systems on real spontaneous conversations and found error rates significantly higher than the best published benchmark results. Your pre-deploy ASR test only predicts production if its audio looks like production audio. Build your test set from sampled real calls, not clean read speech. Our golden dataset guide covers how.

The policy row is the one teams get wrong most often. A disclosure miss might happen in 0.5% of calls. By the rule of three, you would need about 600 clean test calls just to bound it near that level. The only reliable control is to evaluate every production call for policy rules, then alert on the rate.

Mapping a voice turn to OpenTelemetry GenAI spans

Observability needs a data model. For LLM systems, that model is now the OpenTelemetry GenAI semantic conventions. Three facts about them matter before you build on them.

First, they moved. The GenAI conventions now live in a dedicated repository, open-telemetry/semantic-conventions-genai. The old pages on opentelemetry.io say the content has moved and is no longer maintained there.

Second, they are not stable. Every GenAI span and attribute carries Development status. Names have changed before and will change again. Pin a semantic-conventions version in your instrumentation and expect migrations.

Third, they do not define STT or TTS spans. The inference span covers model calls. There are spans for embeddings, retrieval, memory, tools and agents. There is no standard span for a speech recognizer or a speech synthesizer. An open proposal, PR #394, would add realtime voice session spans, a VAD-bounded user speech span, and an audio transcript field. During review its span names changed more than once, so do not hard-code them yet.

The conventions you can use today

ConventionWhat it isUse in a voice agent
Inference span, named `{gen_ai.operation.name} {gen_ai.request.model}`One model callThe LLM step of each turn
`gen_ai.operation.name`, `gen_ai.provider.name`Required on inference spansFilter by provider and operation
`gen_ai.conversation.id`Correlates a conversationSet to your call ID
`gen_ai.prompt.name`, `gen_ai.prompt.version`Named prompt template and versionGroup quality scores by prompt version
`gen_ai.response.time_to_first_chunk`Seconds to first streamed chunkLLM share of turn latency
`execute_tool` operation with `gen_ai.tool.name`, `gen_ai.tool.call.id`, `gen_ai.tool.call.arguments`, `gen_ai.tool.call.result`One tool callBooking, lookup, transfer
`invoke_agent {gen_ai.agent.name}`An agent invocationThe call or the sub-agent
`gen_ai.evaluation.result` eventA score for a GenAI outputAttach evaluator output to the turn it judged

The events document defines `gen_ai.evaluation.result` with `gen_ai.evaluation.name` (required), `gen_ai.evaluation.score.value`, `gen_ai.evaluation.score.label`, and `gen_ai.evaluation.explanation`. It should be parented to the span being evaluated, or carry `gen_ai.response.id` when no span ID is available. This one event is what joins evaluation to observability in a standard way.

The metrics side has `gen_ai.client.operation.duration`, `gen_ai.client.operation.time_to_first_chunk`, `gen_ai.client.token.usage` and `gen_ai.execute_tool.duration`. Those are the series your monitoring layer charts.

One voice turn as a trace

Here is a single turn of an appointment agent, written as a span tree. Durations are illustrative. Custom attributes for the speech layers use a `voice.` prefix, because the spec does not define them.

invoke_agent scheduling-agent        gen_ai.conversation.id=call_8f2c
└── turn 4                           voice.turn.v2v_latency_ms=1180  voice.turn.interrupted=false
    ├── stt                          voice.stt.model=nova-3  voice.stt.confidence=0.71  (custom)
    ├── chat gpt-4.1-mini            gen_ai.prompt.version=sched-v14  time_to_first_chunk=0.41
    │   └── execute_tool reschedule  gen_ai.tool.call.arguments={"day":"thursday"}
    ├── tts                          voice.tts.ttfb_ms=190  voice.tts.chars=58  (custom)
    └── event gen_ai.evaluation.result
          gen_ai.evaluation.name=date_matches_request  score.label=fail
          explanation="Caller asked for Tuesday; tool booked Thursday"

Read it bottom up and the bug explains itself. The evaluator flagged a date mismatch. The tool arguments say Thursday. STT confidence on the turn was 0.71, below this agent's usual range. Pull the audio clip and you will likely hear "Tuesday" spoken over noise. Monitoring would have shown a small rise in date-mismatch failures. Observability shows this exact call and why.

Waterfall of one voice agent turn as OpenTelemetry spans: STT, LLM inference, tool call and TTS under a turn span, with an evaluation result event attached to the turn it scored

What Pipecat and LiveKit emit, and their gotchas

You rarely build this trace by hand. Both major open-source frameworks emit spans already.

Pipecat's OpenTelemetry tracing organizes traces as conversation, then turn, then stt, llm and tts spans. Turn spans carry `turn.number`, `turn.duration_seconds` and `turn.was_interrupted`. Service spans carry `metrics.ttfb`. STT and TTS spans reuse `gen_ai.provider.name` and `gen_ai.request.model`, even though the spec defines those for model calls. Two defaults bite people. `enable_turn_tracking` defaults to `False` on the `PipelineWorker`, so you get no turn spans unless you set it. And the docs note that missing service data usually means `enable_metrics=True` was not set in `PipelineParams`.

LiveKit Agents exports its session spans once you call `set_tracer_provider` from `livekit.agents.telemetry`. Attributes come from an `lk.` namespace and the `gen_ai.` namespace. Three details are worth knowing:

  • Agents 1.7.0 renamed 12 content attributes with an `lk.pii.` prefix, so `lk.chat_ctx` became `lk.pii.chat_ctx`. LiveKit's docs warn that dashboards and queries using old names stop matching without raising an error. That is a monitoring failure caused by an upgrade, and only a test of the dashboards would catch it.
  • Passing `allow_pii=False` strips content attributes, tool arguments and results, and exception messages before any exporter sees them. Token counts, model names and durations remain.
  • `gen_ai.usage.input_tokens` already includes `gen_ai.usage.cache_read.input_tokens`. Adding the two double-counts cached tokens and inflates cost dashboards.

Two gotchas when attaching evaluation to traces

Most evaluation runs after the call ends. By then the turn span is closed, and you cannot add events to an ended span. Store the trace ID and span ID of each turn when the call runs. The evaluator then creates its own span in the same trace, parented to the stored context.

# Simplified, illustrative. Uses the OpenTelemetry Python API.
from opentelemetry import trace
from opentelemetry.trace import NonRecordingSpan, SpanContext, TraceFlags

tracer = trace.get_tracer("voice-evaluator")

def emit_eval(trace_id: int, span_id: int, name: str,
              score: float, label: str, explanation: str) -> None:
    turn_ctx = SpanContext(
        trace_id=trace_id, span_id=span_id, is_remote=True,
        trace_flags=TraceFlags(TraceFlags.SAMPLED),
    )
    parent = trace.set_span_in_context(NonRecordingSpan(turn_ctx))
    with tracer.start_as_current_span(f"evaluate {name}", context=parent) as span:
        span.add_event("gen_ai.evaluation.result", attributes={
            "gen_ai.evaluation.name": name,
            "gen_ai.evaluation.score.value": score,
            "gen_ai.evaluation.score.label": label,
            "gen_ai.evaluation.explanation": explanation,
        })

The convention defines `gen_ai.evaluation.result` as a log-based event. Some SDKs do not emit those yet, so this sketch records it as a span event. The second gotcha is sampling. If your collector uses tail sampling, it decides to keep or drop each trace after a fixed wait. A score that arrives minutes later may land in a trace the collector already dropped. Either keep every trace that will be scored, or write scores to your call store keyed by call ID as well.

Finally, keep audio out of spans. PR #394 reviewers noted that a minute of raw audio is megabytes. Store clips in object storage and put the URI on the span. Our what to log on every call guide lists the full field set.

How many calls to score: the sampling math

Once evaluation runs on production traffic, the next question is how much of it to score. The answer depends on the smallest change you need to detect.

Detecting a regression between two versions

To compare a failure rate between a baseline (p1) and a new version (p2), the standard two-proportion sample size per arm, at 5% two-sided significance and 80% power, is:

n = (1.96 × sqrt(2 × p̄ × (1 − p̄)) + 0.84 × sqrt(p1(1 − p1) + p2(1 − p2)))² / (p1 − p2)²

where p̄ is the average of p1 and p2. Plugging in:

Baseline failure rateNew failure rateScored calls needed per arm
2%3%3,826
4%5%6,745
5%6%8,158
5%8%1,059
5%10%435

Two lessons come out of this table. A 200-scenario test suite can catch a doubling of failures, but it cannot see a one-point rise. And a one-point rise needs thousands of scored calls on each side.

The worked example

Assume an agent takes 10,000 calls a day and you want to catch a task-failure rise from 4% to 5% after a prompt change.

  • Scoring a 2% random sample gives 200 scored calls a day. Reaching 6,745 per arm takes about 34 days. The regression runs for a month before the data can show it.
  • Scoring every call gives 10,000 scored calls a day. You reach 6,745 per arm in well under a day on a 50/50 canary.

This is the core argument behind scoring every call, not a sample. Sampling is not wrong. It just buys you slow detection.

Reading a single rate: use Wilson intervals

For a dashboard tile showing one failure rate, show the interval, not just the point. The Wilson score interval behaves well at the small rates voice teams care about. At 95% confidence:

Failures / scored callsPoint estimateWilson 95% interval
5 / 1005.0%2.2% to 11.2%
20 / 4005.0%3.3% to 7.6%
50 / 1,0005.0%3.8% to 6.5%
200 / 4,0005.0%4.4% to 5.7%

A daily tile built on 100 scored calls can swing from 2% to 11% on noise alone. Teams chase those swings and lose trust in the dashboard. Either score more calls or widen the time window.

What scoring every call costs

Cost is the usual objection to full coverage. Work it through with explicit assumptions.

Assume 300,000 calls a month, 4 minutes each. Assume the evaluator reads about 3,000 input tokens per call (transcript, tool log, rubric) and writes 300 output tokens. That is 900 million input tokens and 90 million output tokens a month.

  • At an illustrative LLM price of $1.00 per million input tokens and $4.00 per million output tokens, scoring costs about $900 + $360 = $1,260 a month.
  • The voice minutes cost far more. 1.2 million minutes at GPT-Live's $0.05 per minute is $60,000 a month.

Under these assumptions, scoring every call adds about 2% to the voice bill. If the judging is mostly classification, a dedicated classifier lowers it further. Jev from TypeSafe System One is one example: highly consistent, priced at $0.042 per million input tokens, with 70 to 500 ms responses. At 900 million input tokens that is about $38 a month.

Where you should sample is human review. People are expensive and slow. A sound pattern: humans review every call the evaluator flags as a severe failure, plus a small random sample each day. The random sample measures how often the evaluator agrees with humans. Our guide to sampling live calls covers stratification.

Your rubric will drift, and that is expected

Human review does more than check the evaluator. It changes what you evaluate. In Who Validates the Validators?, Shankar and colleagues built a tool for aligning LLM graders with human grades. They observed what they called criteria drift: people need criteria to grade outputs, but grading outputs changes their criteria. Some criteria could only be defined after seeing real outputs.

For voice teams this is the formal version of a familiar experience. You read 50 real calls and discover a failure you never wrote a rule for. That is observability feeding evaluation. Version your rubric, re-score a fixed reference set when it changes, and never compare scores across rubric versions without that re-score. Our post on LLM judge limits covers the judge side.

Alert design: SLO burn rates for a voice agent

Monitoring is only useful if it pages the right person at the right time. The SRE Workbook chapter on alerting on SLOs gives a tested method. Voice agents need a few changes to it.

Define SLIs as good events over total events

The workbook recommends SLIs that count good events against total events. Burn rate is how fast you spend the error budget relative to the SLO. A burn rate of 1 uses exactly the whole budget over the SLO window.

Percentile targets like "p95 under 1,200 ms" do not fit this model well. You cannot average percentiles across agent servers. Convert them into ratios instead: "95% of turns have voice-to-voice latency at or under 1,200 ms." Then count slow turns as bad events from a latency histogram.

Here are three voice SLIs, with illustrative targets:

SLIGood eventIllustrative SLO (30 days)Error budget
Turn latencyAgent turn with voice-to-voice latency at or under 1,500 ms99%1% of turns
Task successScored call where the evaluator marks the task complete and correct95%5% of calls
Audio pathCall with two-way audio and no transport drop99.5%0.5% of calls

Voice-to-voice latency is the gap between the end of caller speech and the first agent audio frame sent back. Measure it at the transport, not inside the LLM span.

Apply the multiwindow, multi-burn-rate thresholds

The workbook's recommended starting point is three tiers. Each needs both a long and a short window to exceed the threshold, so alerts stop soon after a problem ends.

SeverityLong windowShort windowBurn rateBudget consumed
Page1 hour5 minutes14.42%
Page6 hours30 minutes65%
Ticket3 days6 hours110%

Translate that into failure-rate thresholds for each voice SLI:

SLIFast page (14.4x)Slow page (6x)Ticket (1x)
Turn latency, 99% SLOMore than 14.4% slow turnsMore than 6% slow turnsMore than 1% slow turns
Task success, 95% SLOMore than 72% failed callsMore than 30% failed callsMore than 5% failed calls
Audio path, 99.5% SLOMore than 7.2% bad callsMore than 3% bad callsMore than 0.5% bad calls

Look at the task-success row. A 72% failure rate means the agent is broken outright. The workbook warns about this case: with a loose SLO, the fast page tier barely ever fires. For quality SLIs, treat the 6x and 1x tiers as the real alerts. Treat the fast tier as a catastrophe detector.

Here is the latency page rule, adapted from the workbook's Prometheus example:

# Simplified. Assumes a histogram voice_turn_latency_seconds with a 1.5 s bucket.
groups:
- name: voice-turn-latency-slo
  rules:
  - record: voice:slow_turns:ratio_rate5m
    expr: |
      (sum(rate(voice_turn_latency_seconds_count[5m]))
        - sum(rate(voice_turn_latency_seconds_bucket{le="1.5"}[5m])))
      / sum(rate(voice_turn_latency_seconds_count[5m]))
  # Repeat the recording rule for 30m, 1h and 6h windows.
  - alert: VoiceTurnLatencyBudgetBurn
    expr: |
      (voice:slow_turns:ratio_rate1h > (14.4 * 0.01)
        and voice:slow_turns:ratio_rate5m > (14.4 * 0.01))
      or
      (voice:slow_turns:ratio_rate6h > (6 * 0.01)
        and voice:slow_turns:ratio_rate30m > (6 * 0.01))
    labels:
      severity: page

Low traffic and sampled scores change the math

Many voice agents are low-traffic services by the workbook's definition. At night, a 5-minute window may hold three calls. One bad call is then a 33% failure rate and a false page. The workbook offers three fixes: generate artificial traffic, combine related services into one SLI, or lengthen the windows.

Artificial traffic is where testing and monitoring meet. A synthetic caller that dials the production number every few minutes is a black-box prober. It keeps the SLI fed overnight, and it is the same asset your test suite already uses. The workbook also warns about the trap: if real callers hit a failure the synthetic caller does not, the probe's successes mask the real errors. Keep probes in their own series, labeled, and never let them dilute real-caller SLIs.

Sampling also adds noise to quality alerts. Take the task-success ticket rule: alert when the 3-day failure rate tops 5%, with a true healthy rate of 4%. Assume 2,880 calls a day.

Share of calls scoredScored calls in 3 daysChance of a false ticket per check
100%8,640Near 0%
10%864About 6.7%
5%432About 14%

At 5% sampling, roughly one check in seven files a false ticket even when nothing changed. Alert fatigue follows. This is the operational cost of sampling, separate from the slow detection shown earlier.

Alert on voice-specific leading indicators

Burn-rate alerts guard the SLO. A few voice-specific signals deserve dashboards and tickets, not pages:

  • ASR drift. Track the distribution of STT confidence, the rate of agent re-asks such as "could you repeat that," and entity failure rate. Segment each by carrier and codec. Also replay a fixed set of reference audio through your production STT setup every day and compute WER against known transcripts. If WER on unchanged audio moves, the provider changed something. See our post on voice agent metric drift.
  • Prompt regressions. Tag every inference span with `gen_ai.prompt.version`. Roll a new prompt to a canary share. Compare evaluator failure rates by prompt version using the sample-size table above. Our LLM update regression guide covers model swaps the same way.
  • Silent tool failures. Count tool calls that return success with an empty or default payload. They never raise errors. See detecting silent tool failures.

A maturity model for voice agent quality

Use this to place your team and choose the next step. Each level adds one capability and an exit test.

LevelTestingEvaluationMonitoringObservabilityExit test
0. DemoManual calls by the builderGut feelProvider dashboardCall recordings, if enabledCan you name your top three failure modes?
1. ScriptedScenario suite run before releaseRule checks on test transcriptsUptime and error-rate alertsTranscripts searchable by call IDEvery release runs the suite
2. ScoredSynthetic callers with noise, accents, barge-in; k runs per scenarioOne rubric scoring test callsLatency SLI with burn-rate alertsTraces with STT, LLM, tool and TTS spansTest and production use the same scorer
3. Closed loopEvery production failure becomes a regression scenarioSame rubric scores every production callQuality SLIs with ticket and page tiersWide events grouped by prompt, model, carrierA canary can prove a 1-point change
4. AuditedIndependent scenario sets and vendor bake-offsHuman-agreement checks and rubric versioningSynthetic probes kept separate from real-caller SLIsAudio linked to every scored turnAn outside party can reproduce your scores

Most teams that think they have "monitoring" are at Level 1. The jump to Level 2 is the shared evaluator. The jump to Level 3 is volume: scoring enough production calls for the statistics to work.

How to wire testing, evaluation, monitoring and observability together

This runbook takes a team from Level 1 to Level 3. Each step produces an artifact the next step uses.

Loop diagram wiring testing, evaluation, monitoring and observability around one shared evaluator, showing the artifacts passed between them from release gate to production alert to new regression test

1. Write the rubric first. List five to ten pass or fail checks tied to business outcomes: task completed, correct entity captured, required disclosure spoken, correct escalation. Give each check an ID and a version. This rubric is your evaluator's contract.

2. Build one evaluator service. It takes a call bundle (audio URI, transcript, tool calls, metadata) and returns per-check labels, scores and explanations. Test runs and production calls both call it. Never fork it.

3. Instrument the call path. Emit OpenTelemetry spans for the turn, STT, inference, tools and TTS. Set `gen_ai.conversation.id` to the call ID and `gen_ai.prompt.version` on every inference span. In Pipecat, turn on `enable_turn_tracking`. In LiveKit, set the tracer provider and decide on `allow_pii`.

4. Write one wide event per call. Include call ID, agent version, prompt version, model IDs, STT and TTS models, carrier, codec, region, turn count, interruptions, tool errors, latency percentiles and every evaluator label. This row is what you group by when something breaks.

5. Build the pre-release gate. Run the scenario suite with k runs per critical scenario through the evaluator. Block the release if any critical check fails in any run, or if aggregate pass rates fall outside thresholds set from the sample-size table. The voice agent testing guide covers scenario design.

6. Score production traffic. Send every call bundle to the evaluator after hangup. Emit `gen_ai.evaluation.result` linked to the turn span, and also write labels to the wide event, so late scores survive trace sampling.

7. Define SLIs and alerts. Start with turn latency, task success and audio path. Apply the burn-rate tiers. For quality SLIs, rely on the 6x and 1x tiers and label synthetic probes separately.

8. Roll out behind a canary. Ship each prompt or model change to a traffic share large enough to reach the needed sample size within a day. Compare evaluator rates by version before going to 100%. Shadow testing works for changes you cannot expose to callers.

9. Triage with traces. When an alert fires, filter wide events by the failing check, group by high-cardinality fields, and open the traces and audio for the top cluster. Write down the cause.

10. Close the loop. Convert each confirmed production failure into a regression scenario with the caller audio pattern that triggered it. If the cause was a missed criterion, add a rubric check, bump the rubric version, and re-score a fixed reference set. The production feedback loop post covers the cadence.

Where independent evaluation fits

The weak point in this whole system is the evaluator. If the team that builds the agent also writes and tunes the scorer, the scorer tends to drift toward what the agent already does well. That is not bad faith. It is criteria drift, seen from inside one team.

Evalgent is an independent evaluation platform for voice agents. It covers the parts of this runbook where a neutral scorer matters most: the pre-launch audit, vendor bake-offs on identical scenarios, regression suites on every release, and scoring production calls against the same rubric used before launch. It does not replace your tracing backend or your alerting stack. It feeds them consistent scores. For a wider view of how these failures appear in the field, see why voice agents fail in production and our production monitoring guide.

Frequently asked questions

What is the difference between voice agent monitoring and observability?

Monitoring watches known indicators, such as turn latency or task-failure rate, and alerts when one crosses a threshold. Observability lets you investigate a specific call and ask questions you did not plan for, using raw traces, audio and high-cardinality fields like prompt version or carrier. Monitoring tells you something is wrong. Observability tells you why.

How do you monitor a voice agent for prompt regressions?

Tag every inference span with `gen_ai.prompt.version`, roll the new prompt to a canary share of traffic, and score every canary and baseline call with the same evaluator. Compare failure rates by version. To detect a rise from 4% to 5% at 80% power, you need about 6,745 scored calls per arm.

How do you detect ASR drift in production?

Track STT confidence distributions, agent re-ask rates and entity failure rates, segmented by carrier and codec. Also replay a fixed set of reference audio through your production STT configuration daily and compute WER against known transcripts. If WER moves on unchanged audio, the provider or your configuration changed.

Is LLM observability enough for voice AI agents?

No. LLM observability covers model calls, prompts and tools. Voice agents also fail in STT, endpointing, TTS and telephony, which LLM traces do not capture. The OpenTelemetry GenAI conventions do not yet define STT or TTS spans. You need audio, VAD timings and transport metrics alongside the LLM spans.

How many production voice agent calls should you score?

Score every call automatically if you need to detect small regressions quickly. A 2% sample of 10,000 daily calls takes about a month to show a one-point change. Under typical assumptions, automated scoring costs a few percent of voice minutes. Sample only for human review, which is slow and expensive.

Do voice agent test assertions carry over to production monitoring?

They do if the same evaluator runs both. Write each assertion as a rubric check with an ID, run it on test calls before release, then run it on every production call after. The pass rate of each check becomes a monitorable SLI, and any production failure can be replayed as a test.

Does OpenTelemetry support voice agents?

Partly. The GenAI semantic conventions cover inference, tool and agent spans, plus a `gen_ai.evaluation.result` event, all at Development status. There are no standard STT or TTS spans yet, and a realtime voice proposal is still in review. Pipecat and LiveKit emit OpenTelemetry traces with their own speech attributes.

What should you look for in a voice agent monitoring tool?

Check five things. Can it hold audio linked to each turn? Can it group by high-cardinality fields like call ID and prompt version? Does it accept OpenTelemetry traces? Can it compute ratio SLIs with burn-rate alerts? Does it score calls with the same rubric your pre-release tests use?

The bottom line

Testing proves a build against cases you chose, monitoring watches known indicators, observability explains the calls nobody predicted, and evaluation is the shared scoring function that lets all three see semantic failures. Build one evaluator, score every call, alert on burn rates, and turn every production failure into a test, and the four practices become one system that improves with each release.

Related Articles