Evalgent
Back to Blog
Voice AI Testing

Testing speech-to-speech voice agents: ground truth, audio-native failures and a test matrix

Deepesh Jayal
Updated
25 min read
Testing speech-to-speech voice agents: ground truth, audio-native failures and a test matrix
On this page

Most guides on this topic stop at "test outcomes, because there is no transcript." That advice is half right. There usually is a transcript. The problem is that it is not evidence of what the model heard or what it said.

This guide covers what vendor docs bury in footnotes and what research papers measured. You get a definition of ground truth for speech-to-speech (S2S) agents and a failure taxonomy that only exists in audio. You also get benchmark numbers to copy and a test matrix with cost and sample-size math. GPT-Live and xAI's Grok Voice are the worked examples.

Everything here is as of October 2, 2026. Model names, prices and event names change. Check the linked docs before you build.

Why speech-to-speech breaks cascaded testing

A cascaded agent has two text contracts. The speech-to-text (STT) output is exactly what the LLM reads. The LLM output is exactly what the text-to-speech (TTS) engine speaks. If the transcript says "fifteen," the LLM saw "fifteen." That is why per-stage testing works on cascades. Our cascading vs speech-to-speech comparison covers the architecture side.

An S2S model breaks both contracts. It consumes audio tokens directly and emits audio tokens directly. Any text you receive is produced alongside the real computation, not as part of it.

OpenAI says this plainly in the Realtime API reference. Input transcription "runs asynchronously" through a separate transcription endpoint, and it "should be treated as guidance of input audio content rather than precisely what the model heard." That line is in the Realtime client events reference. The transcript is produced by a different model (such as `gpt-4o-transcribe`) than the one deciding what to do.

OpenAI's own cookbook shows the gap with a real example. In its out-of-band transcription notebook, the separate ASR model wrote "I had otoo accident." The realtime model understood correctly and replied about an auto accident. The transcript was wrong. The agent was right. A transcript-based grader would have failed a correct call.

xAI's Grok Voice works the same way. Its input transcript events only fire when you set `audio.input.transcription.model` to `"grok-transcribe"`. The `conversation.item.input_audio_transcription.updated` event carries a cumulative transcript that "may include corrections to previous updates," per the xAI speech-to-speech docs. So the caller-side text can change after you have already scored it.

The output side has its own gap. xAI documents a `replace` map that swaps phrases before speech, so "only the spoken audio changes; the transcript the user sees keeps the original text." That is a documented, intentional case where the agent's transcript and the agent's audio disagree.

GPT-Live is the newest design. OpenAI's API launch post says GPT-Live-1 "natively provides ASR transcripts and response text." Even so, the GPT-Live conversations guide says transcript fragments carry "approximate fragment timing rather than exact word boundaries." It also says the context window includes "audio tokens that don't appear in the transcript." For a required disclosure, OpenAI tells you to "use the audio itself to check the wording."

Research models show why the text and audio streams are different objects. Kyutai's Moshi paper introduced "Inner Monologue." The model predicts time-aligned text tokens of its own speech as a prefix to its audio tokens. About 65% of those text tokens are padding. The text stream helps a lot: spoken QA on LlaMA Questions rose from 21.0 to 62.3 with it. But the user's stream is not transcribed at all, and nothing forces the audio to match the text word for word.

HopCascaded agentSpeech-to-speech agent
What the model heardSTT transcript, exactlyAudio tokens; transcript is a guess by a second model
What the model decidedLLM text and tool callsTool calls only; reasoning is not exposed as text
What the caller heardTTS of the LLM textAudio tokens; output transcript is a parallel stream
Where ground truth livesEvery stage boundaryYour script, your app state and the recorded audio

Side-channel transcript: text that a speech-to-speech system emits alongside its audio. It is useful for search, captions and quick checks. It is not proof of what the model heard or said.

The transcript triangle: four texts that should agree

Once you treat transcripts as side-channels, testing gets more precise. Every call has four texts, and each disagreeing pair points at a different broken part.

1. Script (S). What your simulated caller actually said. In a test, you know this exactly.

2. Model input transcript (T-in). What the side-channel ASR thinks the caller said.

3. Model output transcript (T-out). What the model reports it said.

4. Spoken output (A-out). An independent STT pass over the agent's recorded audio.

Add the tool arguments and application state, and you can localize most failures without seeing inside the model.

The speech-to-speech transcript triangle: script, model input transcript, model output transcript and independent STT of agent audio, with tool arguments in the middle and the failure each disagreement reveals
DisagreementWhat it usually meansIs it an agent failure?
Script vs T-inSide-channel ASR misheardOften no. Check tool args before failing the call
Script vs tool argsThe model misunderstood the callerYes. This is the expensive one
Tool result vs A-outSpoken hallucination or dropped readbackYes. The caller heard wrong facts
T-out vs A-outThe transcript does not match the speechYes for compliance; it also corrupts your logs
T-in vs tool args, script agrees with argsTranscript wrong, model rightNo. Fix your grader, not your prompt

The last row matters most. Teams that grade on T-in log false failures like the "otoo accident" case, then "fix" prompts that were fine.

OpenAI's GPT-Live evaluation cookbook uses the same logic in its triage list. If the transcript is right but tool arguments are wrong, inspect the delegation. If the task succeeded but the user never heard the result, inspect result communication. It also asks you to classify each failure as an assistant failure, an infrastructure or evidence problem, or a grader issue.

Our post on transcript vs audio evaluation covers when a transcript-only grader is good enough. For S2S agents, the short answer is: for intent, sometimes; for anything the caller must hear exactly, never.

Audio-native failure modes

Some failures only exist because the model speaks audio tokens directly. A transcript-first harness misses them by design. Here is the taxonomy, with the mechanism behind each one.

Failure modeWhat the caller hearsWhy it happens in S2SHow to detect it
Spoken hallucinationWrong digits, amounts or names, while T-out may look rightAudio and text streams are generated in parallel, not one from the otherIndependent STT of agent audio vs tool result
Truncation desyncAfter a barge-in, the agent acts as if you heard more, or less, than you didAudio is generated faster than playback; truncation guesses the cut point"Count to 50, interrupt at 5, ask where it stopped" probe
Voice driftThe agent's timbre shifts mid-call or toward the caller's voiceVoice is a learned output, not a fixed TTS voiceSpeaker-embedding similarity to the reference voice per turn
Accent driftNon-native accent in non-English outputAccent is sampled with the audio tokensNative-speaker review on a sample; per-language scoring
Language switchingAgent answers in the wrong language after a code-switchAuto language detection on short or mixed utterancesCode-switching scripts; language ID on agent audio
Non-speech artifactsLaughs, breaths, filler or long hums where words should beModel emits non-lexical audio tokensSpeech-duration vs word-count ratio on A-out
Long-session decayForgets early facts or instructions late in the callContext summarization or engine replacementPlant a fact early, probe it after 15+ minutes
Tool calls during speechTwo answers overlap, or a cancelled task still completesAudio and tool events arrive on separate timelinesSlow tool fakes; interrupt during delegation
Safety transfer gapRefusals that hold in text fail in speechSafety tuning does not fully transfer across modalitiesRun your red-team set as audio, not text

A few of these deserve evidence.

Truncation desync. When a caller interrupts, the client tells the server how much audio actually played, using `conversation.item.truncate` with `audio_end_ms`. OpenAI's reference says this deletes the server-side transcript so there is no text in context "that hasn't been heard by the user." Developers on the OpenAI community forum have reported since late 2024 that it does not line up. One reproducible probe from that thread: ask the agent to count to 50, interrupt at 5, then ask where it stopped. The report says it answered "in the 20's or 30's." A March 2026 reply on gpt-realtime-1.5 reported the opposite: the whole turn seemed wiped. Either way, the model's belief about what the caller heard is wrong. That is a community report, not a benchmark, but the probe takes two minutes. Run it on every model you evaluate.

Voice drift and mimicry. OpenAI's GPT-4o system card reported "rare instances" where the model unintentionally emulated the user's voice. The mitigation was a streaming output classifier with precision 0.96 and recall 1.0 in English, and 0.95 and 1.0 in other languages. Red teamers also observed non-native accents in non-English audio output. Your vendor may run a similar classifier. You still want your own per-turn voice similarity score, because a classifier tuned to block extreme cases will not flag gradual drift.

Safety transfer gap. VoiceBench measured LLaMA-Omni refusing 98.46% of harmful requests in text but only 11.35% when the same requests were spoken. That is an open research model, not a commercial agent. The lesson still holds: a guardrail test set run as text proves little about a voice model.

Long sessions. GPT-Live's default context holds 128,000 tokens. When usage passes 90%, it starts a replacement voice engine that receives your instructions and up to 8,192 tokens of history, per the conversations guide. Older details "may be summarized or omitted." Full-Duplex-Bench-v2 found that systems lose track of entities across multi-turn dialogue. Test long calls on purpose. A 3-minute suite will never trigger the summarization path.

For the general version of this problem, see our guide on hallucination rate in voice agents.

What the research says, with numbers you can use

Public benchmarks will not tell you whether your agent works. They tell you where every system tends to fail, which is where to aim your tests.

85% vs 26-51%
Text baseline vs voice agents on tau-voice tasks (arXiv 2603.13686)
0.176-0.588
Self-correction Pass@1 range across systems (Full-Duplex-Bench-v3)
98.46% to 11.35%
LLaMA-Omni refusal rate, text vs speech input (VoiceBench)
92% vs 66%
GPT-4o text vs speech-to-speech on Big Bench Audio at launch

Full-Duplex-Bench: fast is not the same as coherent

Full-Duplex-Bench tests four behaviors: pause handling, backchanneling, smooth turn-taking and user interruption. Its core metric is Takeover Rate (TOR): the share of samples where the model says something other than silence or a short backchannel.

On pause handling, where lower is better, Moshi took over 98.5% of the time on synthetic pauses. Gemini Live took over 25.5% of the time. On user interruption, Moshi stopped and responded every time (TOR 1.000) with 0.257 seconds of latency. But a GPT-4o judge rated the content of those responses 0.765 out of 5. Freeze-Omni was slower at 1.409 seconds and scored 3.615.

The production lesson: score interruption handling on two axes. Did the agent yield, and did what it said next make sense? A latency-only metric would rank the incoherent system first.

Full-Duplex-Bench-v3: self-correction breaks tool calls

Full-Duplex-Bench-v3 used 100 real human recordings from 12 speakers, annotated for fillers, pauses, hesitations, false starts and self-corrections. Six systems ran tool tasks across four domains through LiveKit.

SystemPass@1Turn-take rateTask latencyInterrupt rate
GPT-Realtime0.60096.0%6.89 s13.5%
Gemini Live 3.10.54078.0%4.25 s19.2%
Gemini Live 2.50.49092.0%7.26 s14.1%
Cascaded (Whisper, GPT-4o, TTS)0.450100.0%10.12 s33.0%
Grok0.43094.0%6.65 s25.5%
Ultravox v0.70.41096.0%8.40 s47.9%

Self-correction Pass@1 ranged from 0.176 to 0.588. Hard-tier Pass@1 ranged from 0.200 (Grok) to 0.433 (GPT-Realtime). The authors call self-correction and multi-step reasoning the most consistent failures. If your callers read out account numbers and correct themselves, that is your highest-value test family.

tau-voice: voice keeps 30 to 45% of text capability

tau-voice extended tau-squared-bench to 278 tasks across retail, airline and telecom. It tested OpenAI gpt-realtime-1.5, Google gemini-live-2.5-flash-native-audio and xAI grok-voice-agent in two conditions. Clean used American accents and no noise. Realistic added diverse accents, indoor and outdoor noise, about one noise burst per minute, about 2% frame drops and turn-taking behaviors.

A GPT-5 reasoning baseline in text scored 85%. Voice agents scored 31 to 51% in Clean and 26 to 38% in Realistic. Accents alone cut xAI's retail score from 48% to 30%. The authors traced 79 to 90% of the failures they analyzed to agent behavior, not to the user simulator.

Chart of tau-voice results: text baselines score 85 and 54 percent, while xAI, OpenAI and Google voice agents score 31 to 51 percent clean and 26 to 38 percent realistic

The interaction numbers matter too. OpenAI's agent had 0.90 seconds latency and 100% responsiveness, but only 6% selectivity: it rarely ignored speech it should have ignored. xAI's agent had 57% selectivity but an 84% interrupt rate. These are earlier model versions. The eager-versus-selective trade-off is the one you must tune and test.

VoiceBench: synthetic callers flatter your agent

VoiceBench tested voice assistants on about 6,000 spoken instructions, with speaker, environment and content variations. Three findings change how you build a test set.

  • Every model scored higher on synthetic speech than on real recordings. One model scored about 50% higher on synthetic audio. A TTS-only caller suite overstates quality.
  • Speaking rate has a cliff. End-to-end models degraded badly below 0.5x or above 1.5x speed. A Whisper-based pipeline stayed stable from 0.25x to 2.0x.
  • Mispronunciation hurts most. Average degradation was 20.34% for mispronounced words, 12.55% for repair disfluencies and 2.68% for grammar errors.

Big Bench Audio: the speech reasoning gap

Artificial Analysis built Big Bench Audio from 1,000 Big Bench Hard questions, spoken in 23 synthetic voices. At launch in December 2024, GPT-4o scored 92% text-to-text, 74% text-to-speech and 66% speech-to-speech. Speech in and speech out each cost accuracy.

The field has closed much of that gap. Artificial Analysis's June 2026 Speech-to-Speech Index announcement listed Grok Voice Think Fast 1.0 at 97.1% on Big Bench Audio. The same post put every model below 53% on its tau-voice component. Reasoning on clean questions is close to solved. Agentic tasks over messy audio are not.

BenchmarkWhat it measuresCopy this into your suite
Full-Duplex-BenchPauses, backchannels, turn-taking, interruptionsScore yield and coherence separately
Full-Duplex-Bench-v3Tool use under real disfluencySelf-correction scripts with exact-argument checks
tau-voiceEnd-to-end task success, clean vs realisticSame tasks in clean and degraded audio, report both
VoiceBenchRobustness to speaker, noise, contentReal recordings plus speed and mispronunciation variants
Big Bench AudioReasoning over speechA text-vs-speech control on the same questions

For how these map onto vendor selection, see how to evaluate a speech-to-speech model and our full-duplex voice agents guide.

Building ground truth: verify what the agent actually said

Here is the core technique. You cannot trust the model's transcript of its own speech, so you make your own. Then you compare it to the facts the agent was supposed to say.

Step 1: record two channels on one clock

Record caller audio on the left channel and agent audio on the right. Place both on the session timeline, not on packet arrival time. The harness later in this post shows one way to do it. Stereo makes overlap visible and lets you run STT on each speaker separately.

Step 2: re-transcribe the agent channel independently

Run the agent channel through a different STT model than the one the vendor uses for its side-channel. A Deepgram Nova-3 pre-recorded request with `smart_format=true` returns word-level timestamps, per the Deepgram pre-recorded API reference. Any strong STT works. The point is independence.

Step 3: calibrate the verifier's own error floor

Your verifier makes mistakes too. Measure them before you trust it. Render 100 to 200 known sentences in the agent's voice, pass them through the same codec as production, and transcribe them. The word error rate you get is the verifier's noise floor on this voice and channel. Only flag disagreements above it. Our word error rate guide explains the formula.

Silence is a known hazard. The Careless Whisper study (FAccT 2024) found that roughly 1% of Whisper transcriptions contained entire hallucinated phrases, and 38% of those carried explicit harms. Trim long silences with a VAD before verification, so the verifier does not invent speech the agent never produced.

Step 4: diff on entities, not just words

Word-level disagreement is noisy. A paraphrase like "you're all set" versus "you are all set" is not a failure. What matters is the entities: digits, amounts, dates, names, confirmation codes. Normalize both texts, then check that every value in the tool result appears in the spoken output. Our guide to STT entity accuracy covers normalization rules in depth.

# Simplified ground-truth check for S2S agents. Illustrative, not production.
import os, re, requests, jiwer

DIGITS = {"zero": "0", "oh": "0", "one": "1", "two": "2", "three": "3", "four": "4",
          "five": "5", "six": "6", "seven": "7", "eight": "8", "nine": "9"}

def transcribe_agent_channel(wav_path):
    """Independent STT on the agent-only channel, exported as a mono WAV."""
    with open(wav_path, "rb") as f:
        r = requests.post(
            "https://api.deepgram.com/v1/listen",
            params={"model": "nova-3", "smart_format": "true"},
            headers={"Authorization": f"Token {os.environ['DEEPGRAM_API_KEY']}",
                     "Content-Type": "audio/wav"},
            data=f, timeout=60)
    r.raise_for_status()
    alt = r.json()["results"]["channels"][0]["alternatives"][0]
    return alt["transcript"], alt["words"]  # words carry start/end in seconds

def normalize(text):
    text = re.sub(r"[^\w\s]", " ", text.lower())
    out, run = [], ""
    for tok in (DIGITS.get(t, t) for t in text.split()):
        if tok.isdigit():
            run += tok            # "five five five" and "555" both become "555"
            continue
        if run:
            out.append(run); run = ""
        out.append(tok)
    if run:
        out.append(run)
    return " ".join(out)

def verify_turn(model_transcript, spoken_transcript, tool_values, floor=0.04):
    ref, hyp = normalize(model_transcript), normalize(spoken_transcript)
    disagreement = jiwer.wer(ref, hyp) if ref else 0.0
    not_spoken = [v for v in tool_values if normalize(str(v)) not in hyp]
    return {
        "transcript_vs_speech": round(disagreement, 3),
        "flag_transcript_mismatch": disagreement > floor,  # above verifier noise floor
        "tool_values_not_spoken": not_spoken,              # dropped or hallucinated readback
    }

# Example: tool returned confirmation 48213 for 7:30 p.m.
spoken, words = transcribe_agent_channel("run_017_agent.wav")
print(verify_turn("Your confirmation is 48213 for 7:30.", spoken, ["48213"]))

Step 5: send disagreements to a human ear

Automated checks find candidates. A reviewer listening to a 10-second clip confirms them. Each flagged mismatch is either a real spoken hallucination or a verifier error worth learning from.

Three metrics come out of this process:

  • Spoken-transcript disagreement rate: share of agent turns where T-out and A-out differ beyond the verifier floor.
  • Entity readback accuracy: share of tool-result entities that the agent spoke correctly.
  • Tool-speech consistency: share of calls where the spoken confirmation matches the final application state.

Simulated-caller testing over real audio

Ground truth only helps if the calls look like production. VoiceBench showed that clean TTS callers overstate quality. Build callers that sound like your real ones and use the same audio path.

Pace audio at real time. Stream caller audio in 20 ms frames with a matching sleep. Sending a whole file at once does not simulate a microphone. Keep streaming silence after the caller finishes, because full-duplex models expect a continuous stream.

Use the production codec. Phone calls are 8 kHz G.711. The narrowband channel cuts high frequencies where many consonants live, which is why "fifteen" and "fifty" get confused on phone lines. Both GPT-Live's WebSocket and Grok Voice accept `audio/pcmu` at 8 kHz, so you can test the real format end to end. Convert recordings with ffmpeg rather than sending 24 kHz audio that never touched a phone line.

# Downsample a 24 kHz caller recording to 8 kHz mu-law, as a phone line would
ffmpeg -i caller_24k.wav -ar 8000 -ac 1 -c:a pcm_mulaw caller_8k_ulaw.wav

Mix noise at controlled SNR. Pick signal-to-noise ratio levels and mix them reproducibly. Scale the noise so that 10 times log10 of speech power over noise power equals your target.

import numpy as np

def mix_at_snr(speech, noise, snr_db, seed=0):
    """speech, noise: float32 arrays in [-1, 1] at the same sample rate."""
    rng = np.random.default_rng(seed)
    start = rng.integers(0, max(1, len(noise) - len(speech)))
    noise = np.resize(noise[start:], speech.shape)
    p_speech = np.mean(speech ** 2)
    p_noise = np.mean(noise ** 2) + 1e-12
    scale = np.sqrt(p_speech / (p_noise * 10 ** (snr_db / 10)))
    return np.clip(speech + scale * noise, -1.0, 1.0)

def drop_frames(pcm, rate=8000, frame_ms=20, loss=0.02, seed=0):
    """Zero out 2% of 20 ms frames, close to tau-voice's realistic condition."""
    rng = np.random.default_rng(seed)
    out, n = pcm.copy(), rate * frame_ms // 1000
    for i in range(0, len(out) - n, n):
        if rng.random() < loss:
            out[i:i + n] = 0
    return out

Fix the seed per scenario. If the noise changes between runs, you cannot tell a regression from bad luck.

Borrow the research conditions. tau-voice used about one noise burst per minute and about 2% frame drops. VoiceBench found end-to-end models stable between 0.5x and 1.5x speaking rate. Full-Duplex-Bench-v3 annotated five disfluency types. Those give you defensible starting levels instead of guesses.

Mix real and synthetic voices. Use real recordings for your highest-risk scenarios, and TTS voices for coverage across accents and ages. Report the two groups separately, so a synthetic-only pass rate never hides a real-voice failure. Our guide to synthetic callers for voice agent testing covers persona design, and testing STT under background noise has noise recipes.

Measuring latency from audio onset, not events

Event timestamps lie about latency in three ways.

1. Audio arrives in bursts. The server can send several seconds of audio in a fraction of a second. The arrival time of the first chunk is not when the caller hears it.

2. The first chunk may be silent. Many models start with a few hundred milliseconds of near-silence or breath. The caller experiences the first audible sound, not the first packet.

3. Turn events fire late. With server VAD, the end-of-speech event can only fire after the silence window has elapsed. A metric that starts the clock at that event skips the wait the caller actually sat through.

GPT-Live makes this explicit. On WebSocket, `session.output_audio.delta` has no timing fields and there is no output-audio-done event, per the conversations guide. OpenAI's evaluation harness builds a 400 ms playback reserve and notes that its reported latency "is not a model-only measurement." OpenAI defines response latency as time "from the end of audible caller speech to the first qualifying assistant audio."

So measure on the recording. Find the end of caller speech on the left channel. Find the first sustained audible energy on the right channel. Subtract.

import numpy as np

def first_onset_ms(pcm16: bytes, rate=24000, win_ms=10, thresh_dbfs=-40, hold_ms=30):
    """First point where agent audio stays above -40 dBFS for 30 ms."""
    x = np.frombuffer(pcm16, dtype=np.int16).astype(np.float32) / 32768.0
    win = rate * win_ms // 1000
    n = len(x) // win
    rms = np.sqrt((x[: n * win].reshape(n, win) ** 2).mean(axis=1) + 1e-12)
    db = 20 * np.log10(rms)
    need, run = hold_ms // win_ms, 0
    for i, level in enumerate(db):
        run = run + 1 if level > thresh_dbfs else 0
        if run >= need:
            return (i - need + 1) * win_ms
    return None

Tune the threshold once against a few hand-labeled calls, then freeze it. Changing it between runs moves every latency number.

Illustrative timeline comparing event-based latency with audio-onset latency: caller speech ends, the VAD silence window elapses, audio arrives in a burst, then the caller hears the agent

Here is an illustrative phone-path budget. Every number below is an assumption to replace with your own measurements, except the configured silence window.

HopIllustrative assumptionRunning total
Caller stops speaking0 ms0 ms
Carrier and SIP leg inbound80 ms80 ms
VAD silence window (configured `silence_duration_ms`)700 ms780 ms
Model time to first audio chunk450 ms1,230 ms
Leading near-silence in the first chunk120 ms1,350 ms
Client or bridge jitter buffer60 ms1,410 ms
Carrier leg outbound80 ms1,490 ms

An event-based metric would report 450 ms, the time from the turn event to the first chunk. The caller waits roughly 1,490 ms. The biggest single line is the silence window you configured, which no model upgrade will fix. Our guide to time to first audio breaks down each stage, and turn-taking evaluation covers the cut-in side of the trade-off.

Report P50 and P90 across repeated runs, and report them per audio condition. Latency on clean wideband audio says little about latency on noisy 8 kHz calls.

A test matrix and scorecard you can copy

Put it together as a matrix: scenario families down the side, audio conditions across the top. Run every cell several times.

Scenario familyClean 24 kHz8 kHz G.711Babble 10 dB SNRStreet 5 dB SNR2% frame lossNon-native accent
Entity readback (codes, amounts)RequiredRequiredRequiredRequiredRequiredRequired
Self-correction mid-entityRequiredRequiredRequiredOptionalRequiredRequired
Barge-in during answerRequiredRequiredRequiredRequiredOptionalOptional
Barge-in during tool callRequiredRequiredOptionalOptionalOptionalOptional
Backchannel ("mm-hm")RequiredRequiredRequiredOptionalOptionalOptional
Side talk to another personRequiredOptionalRequiredOptionalOptionalOptional
Long pause mid-sentenceRequiredRequiredOptionalOptionalRequiredOptional
Code-switching callerRequiredRequiredOptionalOptionalOptionalRequired
Truncation probe (count, interrupt, ask)RequiredRequiredOptionalOptionalOptionalOptional
Required disclosure wordingRequiredRequiredOptionalOptionalOptionalOptional
Spoken red-team promptsRequiredRequiredOptionalOptionalOptionalRequired
Long call (15+ minutes, planted fact)RequiredRequiredOptionalOptionalOptionalOptional

Worked cost and volume math

Assume all 12 families across all 6 conditions: 72 cells. Run each cell 5 times: 360 calls. At an assumed average of 2 minutes per call, that is 720 minutes.

  • GPT-Live voice layer: 720 x $0.05 = $36, plus backend model and tool charges.
  • Grok Voice: 720 x $0.08 = $57.60. On a free xAI number, add 720 x $0.01 = $7.20, for $64.80. Prices are from the xAI builder launch post and OpenAI's launch post.

Model cost is rarely the constraint. Review time and concurrency are.

How many runs are enough

Two formulas keep you honest.

Rule of three. If a scenario passes every time in n runs, the 95% upper bound on its true failure rate is about 3 divided by n. Zero failures in 5 runs only proves the failure rate is below about 60%. Zero in 60 runs gets you below 5%. Five repeats find flaky behavior. They do not certify a critical path.

Detecting a regression. To detect a task-success drop from 90% to 85% at 95% confidence and 80% power, you need about 686 calls per arm. The standard two-proportion formula is n = (1.96 x sqrt(2 x p-bar x (1 - p-bar)) + 0.84 x sqrt(p1(1 - p1) + p2(1 - p2)))^2 / (p1 - p2)^2, with p-bar = 0.875. For a drop from 95% to 90%, you need about 435 per arm. Small regressions need big suites, which is why you gate on critical-path cells with many repeats, not on the whole matrix.

The scorecard

The thresholds below are suggested starting points, not industry standards. Set your own from your baseline and business risk.

MetricDefinitionSuggested starting bar
Task successFinal app state matches expected stateAt or above your current baseline, per condition
Entity readback accuracyTool-result entities spoken correctly / entities spoken99% or higher on payment and booking flows
Spoken-transcript disagreementTurns where T-out and A-out differ beyond the verifier floorUnder 2% of agent turns
Self-correction successTool args use the corrected value95% or higher on your scripts
Barge-in yield, coherentAgent stops and the next reply uses the new information90% or higher
False yield on backchannelAgent stops for "mm-hm"Under 10%
Audio-onset latency P90End of caller speech to first audible agent soundSet per channel; track drift release to release
Disclosure verified in audioRequired wording found in A-out, not just T-out100%
Voice similarity floorLowest per-turn speaker similarity to the reference voiceNo turn below your calibrated floor
Language matchAgent language equals caller language per turn99% or higher

Score each metric per condition column, then report the worst column, not the average. An agent that passes clean audio and fails at 5 dB SNR is failing your callers on busy streets.

GPT-Live and Grok Voice: what changes for each

Both are speech-to-speech systems with different testing surfaces.

Testing surfaceGPT-Live (`gpt-live-1`)Grok Voice (`grok-voice-think-fast-2.0`)
Price$0.05 per minute voice layer, backend billed separately$0.08 per minute; free xAI number adds $0.01
ReasoningDelegated to a backend model you chooseBuilt in; `reasoning.effort` defaults to `"high"`
Turn detection controlsNone exposed; the model decides`server_vad` with `threshold` (default 0.85), `silence_duration_ms`, `prefix_padding_ms` (default 333)
Caller transcriptNative ASR transcript fragments with approximate timingOnly with `grok-transcribe`; cumulative updates that can revise earlier text
Known transcript divergenceAudio tokens not in transcript; check disclosures in audio`replace` map changes speech, not transcript
Tool-call gotchaInterrupting speech does not cancel backend workSending `response.create` too early overlaps audio
VersioningOne listed snapshot, `gpt-live-1`Pin a version; `grok-voice-latest` moves

GPT-Live. The split between voice layer and backend means two things to test: did the voice layer hear and delegate correctly, and did the backend act correctly? Session forking lets you replay the same conversation point across variants, which is the cleanest regression tool any S2S API offers today. Our GPT-Live voice agent testing guide has a full harness, and the build guide covers setup.

Grok Voice. xAI's docs describe a specific overlap bug pattern. When the model calls a tool mid-response, the server sends all audio first, then the function call events with `response.done`. If your client returns the tool result and sends `response.create` right away, the next answer starts while the previous one is still playing. xAI recommends waiting until playback is complete or nearly complete. Test this with a slow tool fake and a long spoken preamble.

Two more Grok details catch teams out. A `language_hint` for Spanish or Portuguese must be regional, like `es-MX` or `pt-BR`. Unrecognized codes are "silently ignored," falling back to auto-detection. And `idle_timeout_ms` re-prompts silent callers repeatedly until they speak, which keeps dead calls billing. For pricing and setup, see our Grok Voice agent guide.

Here is a simplified capture harness for Grok Voice, using event names from the xAI docs. It streams paced caller audio, places agent audio on a playback timeline, and keeps every transcript event for the triangle checks.

# Simplified Grok Voice capture harness. Illustrative, not production.
import asyncio, base64, json, os, time, wave
import websockets

MODEL = "grok-voice-think-fast-2.0"  # pin a version; grok-voice-latest moves
RATE, FRAME_MS = 24000, 20
FRAME_BYTES = RATE * 2 * FRAME_MS // 1000
JITTER_S = 0.06                      # assumed client playback buffer

SESSION = {"type": "session.update", "session": {
    "voice": "eve",
    "instructions": "You are Acme support. Read order numbers back digit by digit.",
    "turn_detection": {"type": "server_vad", "silence_duration_ms": 700},
    "audio": {
        "input": {"format": {"type": "audio/pcm", "rate": RATE},
                  "transcription": {"model": "grok-transcribe"}},
        "output": {"format": {"type": "audio/pcm", "rate": RATE}},
    },
}}

async def run(caller_wav, tail_s=10):
    t0 = time.monotonic()
    agent = bytearray()                      # agent audio on the session timeline
    log = {"t_out": [], "t_in": {}, "tools": [], "caller_end_s": None}

    def place(chunk, recv_s):
        cursor = len(agent) / (2 * RATE)
        start = max(cursor, recv_s + JITTER_S)  # play no earlier than arrival
        agent.extend(b"\x00" * (int((start - cursor) * RATE) * 2))
        agent.extend(chunk)

    url = f"wss://api.x.ai/v1/realtime?model={MODEL}"
    headers = {"Authorization": f"Bearer {os.environ['XAI_API_KEY']}"}
    async with websockets.connect(url, additional_headers=headers) as ws:
        await ws.send(json.dumps(SESSION))

        async def send_caller():
            with wave.open(caller_wav, "rb") as w:
                pcm = w.readframes(w.getnframes())
            frames = [pcm[i:i + FRAME_BYTES] for i in range(0, len(pcm), FRAME_BYTES)]
            frames += [b"\x00" * FRAME_BYTES] * (tail_s * 1000 // FRAME_MS)
            for i, f in enumerate(frames):
                if i * FRAME_BYTES >= len(pcm) and log["caller_end_s"] is None:
                    log["caller_end_s"] = time.monotonic() - t0
                await ws.send(json.dumps({"type": "input_audio_buffer.append",
                                          "audio": base64.b64encode(f).decode()}))
                await asyncio.sleep(FRAME_MS / 1000)
            await ws.close()

        sender = asyncio.create_task(send_caller())
        try:
            async for raw in ws:
                e, now = json.loads(raw), time.monotonic() - t0
                kind = e["type"]
                if kind == "response.output_audio.delta":
                    place(base64.b64decode(e["delta"]), now)
                elif kind == "response.output_audio_transcript.delta":
                    log["t_out"].append({"t": now, "text": e["delta"]})
                elif kind == "conversation.item.input_audio_transcription.updated":
                    log["t_in"][e.get("item_id")] = e   # cumulative; keep the latest
                elif kind == "response.function_call_arguments.done":
                    log["tools"].append({"t": now, "name": e["name"],
                                         "args": json.loads(e["arguments"])})
        except websockets.ConnectionClosed:
            pass
        sender.cancel()
    return bytes(agent), log

Write the agent bytes and caller audio into one stereo WAV. Then run `first_onset_ms` and the verification step on the agent channel. Tool results are omitted for brevity. In a real harness, return them with `conversation.item.create` and wait for playback before `response.create`. This client also never stops playback on barge-in, so use a fuller client for interruption tests.

How to test a speech-to-speech voice agent

Here is the full procedure in order. It works for GPT-Live, Grok Voice, OpenAI Realtime models and other S2S APIs.

1. Write scenarios with hidden expectations. For each one, define the caller script, the expected tool calls with exact arguments, the expected final state and any wording the caller must hear.

2. Build fake backends with resettable state. Every run starts from the same seeded data, so differences come from the agent, not the fixture.

3. Stream paced, production-shaped audio. Send 20 ms frames in real time, at your production codec and sample rate, with a continuous silent tail.

4. Record two channels on the session timeline. Caller left, agent right, with agent chunks placed at playback time rather than arrival time.

5. Log every side-channel text with timestamps. Keep input transcripts, output transcripts and tool events, including revisions to cumulative transcripts.

6. Calibrate an independent verifier. Measure its error floor on the agent's voice through your codec, then re-transcribe every agent turn.

7. Score the triangle. Check tool args against the script, spoken entities against tool results, and the output transcript against the spoken audio.

8. Measure latency from audio onset. Use the end of caller speech and the first sustained audible agent sound. Report P50 and P90 per condition.

9. Run the matrix with repeats. Use at least 5 runs per cell for discovery, and many more on critical-path cells you gate releases on.

10. Probe the audio-native failures directly. Run the truncation probe, code-switching, long calls with planted facts, spoken red-team prompts and slow-tool overlap tests.

11. Triage before you fix. Label each failure as agent, infrastructure or grader. Fix grader errors first, because they corrupt every other number.

12. Re-run on every change. Prompt edits, backend swaps, voice changes and silent model updates can all regress an S2S agent. Our guide to LLM update regressions explains why.

For the wider discipline behind this procedure, see our AI voice agent testing pillar.

Where independent evaluation helps

An S2S vendor's dashboard grades calls with that vendor's own transcripts. That is the side-channel this whole guide warns about. An independent evaluator brings its own verifier, its own callers and its own clock.

Evalgent is that independent layer. It runs simulated callers over real codecs and noise against any S2S provider, re-transcribes agent audio independently, and scores tool arguments, spoken entities, latency from audio onset and turn-taking against thresholds you set. Teams use it for pre-launch audits, for bake-offs between GPT-Live, Grok Voice and cascaded stacks on the same scenarios, and for regression runs whenever a model or prompt changes.

Frequently asked questions

How do you test a speech-to-speech voice agent?

Stream realistic caller audio at real-time pace, record both channels on one clock, and log every transcript and tool event. Then check tool arguments against your script, re-transcribe the agent's audio independently to confirm what it said, and measure latency from audible onset. Repeat each scenario across codecs, noise levels and accents, and report the worst condition.

Can you trust the transcript from a realtime voice model?

Not as ground truth. OpenAI's Realtime reference says input transcription runs asynchronously on a separate model and is guidance, not what the model heard. xAI's input transcripts can revise earlier text, and its replace map changes speech without changing the transcript. Use transcripts for search and quick checks, and verify anything critical against the audio.

How do you verify what a voice agent actually said?

Run the agent's audio channel through an independent STT model, normalize numbers and names, and compare it with the tool result and the model's own output transcript. Calibrate the verifier's error rate on the agent's voice first, trim long silences to avoid STT hallucinations, and send flagged mismatches to a human reviewer.

How do you measure speech-to-speech latency correctly?

Measure from the end of caller speech to the first sustained audible agent sound on the recording. Event timestamps mislead, because audio arrives in bursts, first chunks can be silent, and turn events fire after the silence window. Report P50 and P90 per audio condition, and track the configured silence window separately.

How do you test Grok speech to speech?

Connect to wss://api.x.ai/v1/realtime with a pinned model such as grok-voice-think-fast-2.0. Enable grok-transcribe if you need caller transcripts, and keep the latest cumulative update. Test server_vad settings, slow tool calls for audio overlap, regional language hints, and entity readback over 8 kHz audio. Budget $0.08 per minute plus telephony.

What failure modes are unique to speech-to-speech models?

Spoken hallucinations that the transcript hides, truncation desync after barge-in, voice and accent drift, wrong-language replies after code-switching, non-speech artifacts like laughs or hums, long-session memory loss, overlapping audio around tool calls, and safety refusals that hold in text but fail in speech. Each needs an audio-level check.

How many test calls do you need for a voice agent?

It depends on the claim. Zero failures in n runs bounds the true failure rate at about 3 divided by n, so 60 clean runs show under 5%. Detecting a drop from 90% to 85% task success needs about 686 calls per arm. Use small repeats for discovery and large repeats on critical paths.

Do research benchmarks predict production performance?

They predict where systems fail, not how your agent will score. tau-voice found voice agents keep only 30 to 45% of text-agent capability, and Full-Duplex-Bench-v3 found self-correction is a consistent weak spot. Use those findings to design scenarios, then measure on your own callers, tools and phone lines.

The bottom line

Speech-to-speech agents do have transcripts, but those transcripts are side-channels, so real testing means building your own ground truth from scripts, app state and independently verified audio. Test the audio-native failures directly, measure latency from what the caller hears, and gate releases on the worst audio condition, not the average.

Related Articles