Voice Agent Benchmarks Explained: VoiceBench, VoiceAgentBench, Full-Duplex-Bench, τ-Voice and More

On this page
Every voice model launch ships with a benchmark chart, each using a different benchmark, setting or version. None tells you whether the model will book a dental appointment over a Twilio SIP trunk.
This guide goes through each major benchmark: what it measures, its data, its headline result, and what it misses for a US phone agent. Then come three patterns visible only side by side, a decision guide, and a test plan for the gaps.
What a voice agent benchmark measures
A voice agent has four jobs: understand speech, manage the floor (talk, wait, stop, ignore "mm-hmm"), act correctly through tools and policy, and speak audio a caller can follow. Most benchmarks cover one or two. A model that tops a reasoning benchmark can still interrupt every turn, so the first question is which benchmark measures the failure you care about. For model architecture trade-offs, see our guide to evaluating speech-to-speech models.
VoiceBench: spoken instructions under realistic variation
VoiceBench (Chen et al., 2024) was the first broad benchmark for LLM-based voice assistants. It asks one question: does the model answer a spoken instruction as well as the same instruction in text?
Data. The repository lists 11 subsets. Most are Google TTS audio (AlpacaEval, OpenBookQA, MMSU, IFEval, AdvBench, a 46-sample MT-Bench). Four are human speech: CommonEval (200), SD-QA (553 questions in 11 English accents), WildVoice (1,000 crowd-sourced) and a human-read BBH set (1,000). All English.
Metrics. GPT-4o-mini scores open answers from 1 to 5. Multiple-choice and SD-QA use accuracy, IFEval uses its rule checks, and AdvBench uses refusal rate. Only the text of the reply is scored, not the audio.
Headline results. A pipeline of Whisper-large-v3 plus GPT-4o scored 87.23 overall on spoken input. GPT-4o-Audio scored 86.42. Open end-to-end models trailed the simple pipelines by more than 20 points. The paper also ran perturbation studies:
- Speaking rate. End-to-end models degraded below 0.5x or above 1.5x speed. The pipeline held from 0.25x to 2.0x.
- Content noise. Mispronunciation cut scores 20.34% on average across models, the worst of any perturbation. Self-repairs ("Tuesday, no, Wednesday") cut 12.55%.
- Safety transfer. LLaMA-Omni refused 98.46% of harmful requests in text but 11.35% when spoken.
What it misses for US phone agents. No tools, no backend state, almost no multi-turn. The environment tests include packet loss and a low-pass "far-field" filter, but no telephony codec. The accent subset uses short trivia questions, not names, emails or order numbers.
Use it for: a quick check of whether a speech model loses text capability, and for borrowing perturbation recipes.
VoiceAgentBench: tool calls from speech
VoiceAgentBench (Krutrim AI Labs, 2025) asks whether speech models can do agent work: pick tools, fill arguments, chain calls and refuse unsafe ones.
Data. Over 6,000 spoken queries, all synthesized with TTS. The authors pick speaker embeddings to vary the voices. Coverage: English and six Indic languages. Tools come from the Berkeley Function Calling Leaderboard plus custom agents, and many queries are rewritten in an Indian context.
Task types. Single tool, single tool with retrieval, parallel calls, sequentially dependent calls, dialog-based calls (the final call after a multi-turn exchange) and safety refusals.
Metrics. Tool selection, structural consistency, parameter filling (judged by GPT-4o-mini, checked against human agreement) and refusal rate.
Headline results. The best system reached only 60.6% average parameter-filling accuracy in English. Whisper-v3 plus a text LLM (Qwen3-8B, Gemma3-27B or Llama 3.3 70B) beat the 7B end-to-end speech models, and the speech models degraded more on Indic languages. Sequentially dependent calls scored lowest for every model.
What it misses. The models under test are open-weight. No commercial realtime API appears in the main tables. Queries are single utterances from TTS, with no live turn-taking, no latency and no telephony. The entities are Indian-context, not US addresses, Medicare IDs or ZIP codes.
Use it for: evidence that argument filling, not tool choice, is where spoken tool use breaks. Our tool call accuracy guide shows how to measure that on your own schema.
Full-Duplex-Bench: four versions, four different questions
Full-Duplex-Bench is now a family. The repository holds v1, v1.5, v2 and v3, and they answer different questions. Vendors and blogs often cite "Full-Duplex-Bench" without saying which one.
v1: turn-taking behaviors
v1 (Lin et al., ASRU 2025) scores four behaviors: pause handling, backchanneling, smooth turn-taking and user interruption.
- Data. Real two-channel speech from Candor (216 pause samples, 119 turn-taking samples) and ICC (55 backchannel stimuli rated by 118 listeners). Synthetic ChatTTS dialogues supply 200 interruptions and 137 pause samples.
- Metrics. Takeover rate (TOR) during pauses, backchannel frequency and timing divergence (JSD), response latency, and a GPT-4o relevance score after interruptions.
- Headline. On synthetic pauses, Moshi took the floor 98.5% of the time. Gemini Live took it 25.5% of the time. dGSLM and Moshi answered in about 0.3 s but often interrupted.
- Setup detail that matters. Gemini Live was fed 16 kHz PCM in 30 ms chunks, and a new session started after each reply. That is a probe of single-turn behavior, not of long-call stability.
v1.5: overlap handling
v1.5 (ICASSP 2026) plays overlapping speech while the model talks: user interruptions (200), backchannels (99), the user talking to someone else, and background speech (100). It classifies each reaction as respond, resume, uncertain or unknown. It also measures stop latency, response latency, prosody shifts and UTMOSv2.
The headline is a split in strategy. GPT-4o Realtime stopped within 0.18 to 0.23 s of overlap onset. Gemini and Nova Sonic took 2.20 s and 2.25 s to stop for real interruptions. But GPT-4o Realtime also responded to 91% of side conversations and 93% of background speech. Nova Sonic resumed through 98% of background speech. One system yields to everything, another yields to almost nothing.
The background-speech condition is low-pass filtered at 3 kHz, close to phone bandwidth. That filter applies only to the distractor, not to the caller.
v2: live multi-turn with an examiner
v2 runs live sessions through a WebRTC orchestrator at 48 kHz in 10 ms frames. The examiner is itself a spoken model (gpt-realtime with synthesized speech), and it pursues staged goals at Fast or Slow pacing. Task families are Daily, Correction, Entity Tracking and Safety. A Gemini 2.5 Flash judge scores turn-taking fluency, instruction following and a task metric from 1 to 5. Judge agreement with humans was moderate (Pearson r of 0.59 to 0.69).
All systems degraded as conversations lengthened, and instruction following fell faster than turn-taking. GPT-Realtime scored above 4.0 on all three task metrics under Fast pacing. Moshi and Freeze-Omni stayed below 3.0 on Correction and Entity. Note that the examiner and the top evaluatee come from the same model family. The paper validates the judge against humans, but no ablation swaps the examiner.
v3: tool use under real disfluency
v3 is the most relevant to phone agents. It has 100 real recordings from 12 speakers, including non-native speakers with Korean and Russian backgrounds, captured on everyday built-in microphones. Each scenario requires chained API calls across four domains, and 21 scenarios contain a mid-utterance self-correction.
All six systems ran through LiveKit's realtime agent framework. GPT-Realtime led with Pass@1 of 0.600 and a 13.5% interruption rate. On self-correction scenarios it scored 0.588. The cascaded baseline scored 0.176 on self-corrections, 0.450 overall, and had the highest latency at 10.12 s.
That cascaded baseline was Whisper plus GPT-4o plus OpenAI TTS. It is not a tuned streaming pipeline with conversational STT and end-of-turn detection, which is what most LiveKit and Pipecat teams run. Treat it as a floor, not as "cascaded."
What the family misses. No telephony codec in any version. v1 and v1.5 are single exchanges, v2 uses a synthetic examiner, and in v3 one scenario moves Pass@1 by a full point. For how these behaviors show up in production, see our full-duplex voice agents guide and the backchanneling and proactivity post.
HumDial: real human overlap at scale
The ICASSP 2026 HumDial Challenge has two tracks: Emotional Intelligence and Full-Duplex Interaction. The full-duplex study describes the second.
Data. More than 100 hours of human-recorded, dual-channel interactive speech in Chinese and English, with train, validation and test splits. Scripts come from an LLM; people record them.
Scenarios. Interruption covers follow-up questions, negation or dissatisfaction, repetition requests, topic switches and "stop talking." Rejection covers backchannels, pause handling, third-party speech and speech directed to others. Most categories have 200 validation and 600 test samples.
Metrics. Responses are transcribed (Parakeet-TDT for English), then DeepSeek-V3 classifies the behavior following the Full-Duplex-Bench v1.5 protocol. Latency is scored too.
Headline. Among baselines, Gemini scored 79.8 on interruption handling but 36.5 on rejection. Freeze-Omni led rejection at 50.2 but scored 29.6 on interruption. Gemini had the lowest latency at 1.301 s.
What it misses. No tools, tasks or telephony, and half the data is Mandarin. It is still the best public source of real backchannels and side speech for testing interruption logic.
VocalBench: the one that grades the spoken answer
VocalBench (Shanghai Jiao Tong University and partners) is the main benchmark that scores the audio a model produces, not just the text.
Data. About 24,000 instances across 14 capabilities in English and Mandarin. Groups cover semantic quality, acoustic quality, chat ability and robustness. Audio is synthesized with CosyVoice. The authors list human recordings as future work.
Metrics. Accuracy and LLM-judge scores for content. UTMOS and WER for the output speech. Empathy, safety, latency and real-time factor. Robustness reuses VoiceBench's perturbations (noise, reverb, far-field, clipping, packet loss).
Headline. Across 27 systems, a cascade with GPT-4o had the top English overall score (82.68). Qwen3-Omni came next (78.78). Cascades kept a slight edge on content but scored worse on emotional empathy. The authors trace that gap to how the voice sounds, not to what it says. They also found models that emit structured text, such as markdown lists in math answers, which a listener can't follow.
What it misses. Synthetic input only. No tools. Its UTMOS is computed on wideband output, not on what survives an 8 kHz phone leg.
MultiChallenge and Audio MultiChallenge: same name, different tests
This is where vendor charts most often mislead.
MultiChallenge (Scale AI, 2025) is a text benchmark. It tests multi-turn instruction retention, inference memory, self-coherence and versioned editing with instance-level rubrics. At launch, the top model (Claude 3.5 Sonnet) scored 41.4%.
In August 2025, OpenAI reported gpt-realtime at 30.5% on "MultiChallenge audio," up from 20.6%. That version was OpenAI's own TTS conversion of an audio-friendly subset.
Audio MultiChallenge (Scale AI, December 2025) is a separate benchmark built on real speech:
- Data. 452 conversations from 47 speakers, about 15 hours of user audio recorded at 48 kHz with no post-processing. It adds a fourth axis, Voice Editing: mid-utterance repairs and backtracking.
- Metrics. 1,712 instance-specific rubrics. A task passes only if every rubric passes. The rubric judge agreed with humans 93% of the time.
- Headline. Gemini 3 Pro Preview (Thinking) led at 54.65%. GPT Realtime scored 23.45% with text output and 20.35% with audio output. Self-coherence fell from 33.3% on tasks with under 60 s of user audio to 20.0% at 3 to 5 minutes. Recalling cues carried only in the audio, such as tone or a background sound, scored 36.5% lower in relative terms than recalling words.
The paper's own footnote warns that some models "previously report scores on a TTS set of MultiChallenge," which it distinguishes from its own. So when a chart says "MultiChallenge," check which one.
What it misses. No tools, no turn-taking timing, no telephony. It does measure the multi-turn memory decay that hits long calls.
τ-Voice: full tasks, policies, and 8 kHz audio
τ-Voice (Ray, Dhandhania, Barres and Narasimhan, 2026) extends τ²-bench to full-duplex voice. It is the closest public benchmark to a real support call.
Data. 278 tasks across Retail (114), Airline (50) and Telecom (114). Each task has a verifiable database end state and a domain policy. A simulated caller speaks through seven TTS personas (ElevenLabs v3 at 24 kHz): American-accented voices for Clean runs and diverse accents for Realistic runs.
Audio. Every condition, including "Clean," passes caller audio through G.711 μ-law at 8 kHz. "Realistic" adds background noise, about one burst per minute, about 2% frame drops from a Gilbert-Elliott model, muffling, diverse accents, coughs, "hold on" side speech, interruptions and backchannels.
Metrics. Pass@1 against the end state, plus four interaction scores: responsiveness, latency, interrupt rate and selectivity (ignoring backchannels, tics and side speech).
Headline results.
- GPT-5 (reasoning) completed 85% in text. Voice agents reached 31 to 51% on clean audio and 26 to 38% on realistic audio.
- In Retail ablations, accents alone cost xAI 18 points and Google 1 point.
- Domain swings were large. OpenAI scored 71% Clean in Retail but 28% in Telecom. xAI led Telecom at 58%.
- No provider won both conversational axes. OpenAI had 0.90 s latency and 100% responsiveness but 6% selectivity. xAI had the best selectivity (57%) but an 84% interrupt rate.
- Authentication was the main bottleneck. Agents failed to capture spelled names and emails, and one said "I've updated your shipping address" without making a tool call.
What it misses. The tested systems are audio-native APIs only. The paper lists cascaded baselines as future work. Tool calls return instantly. The simulator is "more patient than real users." The caller hears the agent's transcript, not its audio, so agent pronunciation is untested. English TTS only. Artificial Analysis now runs a τ-Voice implementation as part of its Speech to Speech Index, averaged over three trials.
Big Bench Audio: spoken reasoning, now near the ceiling
Big Bench Audio (Artificial Analysis, December 2024) converts 1,000 Big Bench Hard questions into speech: 250 each of formal fallacies, navigation, object counting and web-of-lies logic. They are spoken in 23 synthetic voices.
Metric. An LLM judge marks each spoken answer correct or incorrect.
Headline. At launch, GPT-4o scored 92% in text and 66% speech-to-speech. OpenAI reported gpt-realtime at 82.8% in August 2025. In May 2026, OpenAI said GPT-Realtime-2 at high reasoning effort scored 15.2 points above GPT-Realtime-1.5.
Two gotchas.
- The scoring changed. Per the methodology page, v1.0 scored only non-error responses. From v1.1 (March 2026), non-answers count as wrong, and the judge moved to Claude Sonnet 4.6. Scores from before and after March 2026 are not directly comparable.
- Settings change the score. In the same May 2026 release, the Audio MultiChallenge gain was reported at xhigh effort, while the API default is low. Your agent runs whatever effort you configure, and higher effort adds latency.
What it misses. Four puzzle types, clean synthetic audio, one turn, no tools. It says nothing about taking a card number over a noisy cell call.
Voice agent benchmarks compared

| Benchmark | Measures | Input speech | Size | Languages | Tools | 8 kHz phone audio |
|---|---|---|---|---|---|---|
| VoiceBench | Spoken instruction following, knowledge, safety, robustness | Mostly TTS; 4 human subsets | 11 subsets, ~8k items | English | No | No |
| VoiceAgentBench | Tool selection, argument filling, chaining, refusal | TTS | 6,000+ queries | English + 6 Indic | Yes | No |
| Full-Duplex-Bench v1 | Pauses, backchannels, turn-taking, interruption | Real (Candor, ICC) + TTS | ~730 samples | English | No | No |
| Full-Duplex-Bench v1.5 | Overlap reactions, stop latency, prosody | TTS | 4 overlap scenarios | English | No | No |
| Full-Duplex-Bench v2 | Live multi-turn fluency, corrections, entities, safety | TTS examiner | 4 task families | English | No | No |
| Full-Duplex-Bench v3 | Chained tool calls under disfluency | Real recordings | 100 scenarios | English | Yes | No |
| HumDial (full-duplex) | Interruption and rejection behavior | Real, dual-channel | 100+ hours | Chinese, English | No | No |
| VocalBench | Content, output speech quality, empathy, robustness | TTS | ~24k instances | English, Mandarin | No | No |
| Audio MultiChallenge | Multi-turn memory, retention, coherence, voice edits | Real, 48 kHz | 452 conversations | English | No | No |
| τ-Voice | Grounded task completion, policy, duplex dynamics | TTS personas | 278 tasks | English | Yes | Yes (G.711) |
| Big Bench Audio | Spoken reasoning | TTS, 23 voices | 1,000 questions | English | No | No |
Only one row says yes for phone audio. Only three say yes for tools. None uses your caller population, your entity formats or your policy document.
Three patterns that only show up across benchmarks
Read one paper and you get a ranking. Read all of them and three patterns appear.
1. Responsiveness and restraint trade off, in every benchmark

Three independent teams measured the same tension with different data:
- Full-Duplex-Bench v1.5: GPT-4o Realtime stopped fastest for real interruptions and also responded to 91% of side conversations.
- HumDial: Gemini led interruption handling (79.8) but scored 36.5 on rejection. Freeze-Omni led rejection (50.2) and trailed on interruption (29.6).
- τ-Voice: OpenAI had 100% responsiveness and 6% selectivity. Google ignored backchannels and side speech far more often (54% selectivity) but answered only 69% of turns.
No single "turn-taking score" captures this. You need two numbers: how often the agent yields when it should, and how often it holds when it should. In a cascaded stack both move with your barge-in settings, and tuning one moves the other. A dispatch line full of side talk needs restraint. A collections line where callers dispute amounts needs fast yielding.
2. The "cascaded baseline" in papers is not your stack
VoiceBench's pipeline is Whisper-large-v3 plus an LLM, with no streaming. Full-Duplex-Bench-v3's is Whisper plus GPT-4o plus OpenAI TTS, at 10.12 s latency. τ-Voice has no cascaded baseline at all.
Yet the content results favor pipelines. VoiceBench's pipeline beat open end-to-end models by more than 20 points. VoiceAgentBench's ASR-plus-LLM setups beat speech LMs. VocalBench's GPT-4o cascade topped English overall. Latency and turn-taking results come from untuned cascades, so they say little about a tuned streaming pipeline. If you run Deepgram Flux or Nova-3 with Cartesia or ElevenLabs on LiveKit or Pipecat, no public benchmark has tested your configuration. Our cascading vs speech-to-speech comparison covers the architectural trade-offs.
3. Headline numbers depend on settings, versions and domain
- Settings. GPT-Realtime-2's Big Bench Audio gain is at high effort, its Audio MultiChallenge gain at xhigh, and the API default is low.
- Versions. Big Bench Audio changed how it counts non-answers in March 2026. "MultiChallenge audio" (OpenAI's TTS subset) and "Audio MultiChallenge" (Scale's real-speech set) are different tests.
- Domain. On τ-Voice, the same OpenAI model scored 71% Clean in Retail and 28% in Telecom.
Worked math: how much noise sits in a benchmark score. A pass rate's 95% confidence half-width is about 1.96 × √(p(1−p)/n). At p = 0.5:
- τ-Voice, all 278 tasks: 1.96 × √(0.25/278) = ±5.9 points.
- τ-Voice Airline, 50 tasks: 1.96 × √(0.25/50) = ±13.9 points.
- Full-Duplex-Bench-v3, 100 scenarios: ±9.8 points.
A 3-point gap between two models on one domain is noise. A 40-point gap between domains for one model is signal, and your domain is not any of theirs. Our vendor benchmarking guide covers reading vendor charts with these limits.
Which benchmark answers which question
| Your question | Best public benchmark | Read this number | Caveat |
|---|---|---|---|
| Does this model lose capability when input is speech? | VoiceBench, Big Bench Audio | Text vs speech gap on the same items | Clean audio, mostly TTS, near ceiling for top models |
| Will it interrupt callers or stall on "uh-huh"? | Full-Duplex-Bench v1/v1.5, HumDial | Takeover rate on pauses; respond vs resume on backchannels | Single exchanges; no phone codec |
| Can it handle "Tuesday, no, Wednesday" in a tool call? | Full-Duplex-Bench-v3 | Self-correction Pass@1 | 21 scenarios; weak cascaded baseline |
| Does it keep instructions over a 5-minute call? | Audio MultiChallenge, Full-Duplex-Bench v2 | Instruction retention; IF score over time | No tools; v2 uses a synthetic examiner |
| Can it fill tool arguments from speech? | VoiceAgentBench | Parameter-filling accuracy | Open models only; Indian-context entities |
| Will it finish real tasks under policy on phone audio? | τ-Voice | Pass@1 per domain, Clean vs Realistic | Instant tools; TTS callers; no cascades |
| Does the spoken answer sound clear? | VocalBench | UTMOS, output WER | Wideband output, synthetic input |
| Will it work for my callers, my entities, my policy? | None | Your private benchmark | Build it (below) |
For a bake-off shortlist, read τ-Voice per domain, then Full-Duplex-Bench v1.5 or HumDial, then Audio MultiChallenge. Then test your own calls.
What every public benchmark misses for US phone agents
These gaps decide whether a model that scores well publicly works on your phone line.
8 kHz narrowband audio. SIP legs on Twilio or Telnyx commonly carry G.711 at 8 kHz, so energy above 4 kHz is gone and "f" and "s" blur. Only τ-Voice applies the codec. Our phone audio quality guide covers resampling and codec effects.
Your entities. Benchmarks test trivia, BFCL arguments or τ-bench order IDs. Your agent captures member IDs, dates of birth, street names and emails, and τ-Voice found authentication, where these get spelled out, to be the main bottleneck. See our STT entity accuracy guide.
Tool latency and failure. τ-Voice tools return instantly. Real CRM and scheduling APIs take seconds and time out, and the dead air that follows is untested everywhere.
Your policy. Your disclosures, verification steps and escalation rules are not retail or airline policy, and they are where compliance failures happen.
Real callers. No benchmark uses 8 kHz human callers from your region, on speakerphone, in a car. The τ-Voice authors call their TTS accent results "indicative rather than definitive." Our accent robustness guide shows how to build real coverage.
What the caller hears. Nobody scores whether your TTS reads "$1,250.00" or a confirmation code clearly after μ-law compression.
Your stack. Benchmarks test models. You ship VAD, endpointing, STT, LLM, TTS, telephony and a prompt, and a model's win transfers only if the rest matches.
Re-running a public benchmark on phone-band audio
You can close the codec gap for any benchmark that ships audio. This illustrative snippet pushes VoiceBench's US-accent SD-QA split through 8 kHz μ-law and back, as a SIP leg would. The dataset config, split and column names are from the public Hugging Face dataset.
# Illustrative: phone-band a VoiceBench subset before scoring.
# pip install datasets ; requires ffmpeg on PATH
import os, subprocess, tempfile
from datasets import load_dataset, Audio
ds = load_dataset("hlt-lab/voicebench", "sd-qa", split="usa")
ds = ds.cast_column("audio", Audio(decode=False)) # raw bytes, no decoding
def phone_band(src_bytes: bytes, out_path: str) -> None:
with tempfile.TemporaryDirectory() as d:
src, ulaw = os.path.join(d, "src.wav"), os.path.join(d, "ulaw.wav")
open(src, "wb").write(src_bytes)
# Down to 8 kHz G.711 mu-law, like a PSTN/SIP leg...
subprocess.run(["ffmpeg", "-y", "-loglevel", "error", "-i", src,
"-ar", "8000", "-ac", "1", "-c:a", "pcm_mulaw", ulaw], check=True)
# ...then back to 16 kHz PCM, as your STT or model input would receive it.
subprocess.run(["ffmpeg", "-y", "-loglevel", "error", "-i", ulaw,
"-ar", "16000", "-c:a", "pcm_s16le", out_path], check=True)
os.makedirs("sdqa_phone", exist_ok=True)
for i, row in enumerate(ds):
phone_band(row["audio"]["bytes"], f"sdqa_phone/{i:04d}.wav")
# store row["prompt"] and row["reference"] alongside for scoringScore the wideband and phone-band versions with the same harness. The difference is your codec penalty for that model. Repeat with the other SD-QA accent splits (such as `ind_s` or `phl`) to see whether the penalty grows for some accents.
How to build a private voice agent benchmark that covers the gaps

Public benchmarks tell you which models deserve a trial. A private benchmark tells you which one to ship and catches regressions after you do. Here is a plan sized for a small team.
1. Write 40 scenarios from your own call logs, in five families. Use eight each: identity and entity capture, multi-step tool tasks, policy edge cases, mid-utterance corrections, and off-script behavior (side talk, "hold on," silence). Give each a single verifiable end state, as τ-Voice does, so pass or fail doesn't depend on a judge's taste. Our golden dataset guide covers sourcing and labeling.
2. Run every scenario over a real phone path. Place calls through your actual Twilio or Telnyx trunk, or at minimum apply G.711 μ-law at 8 kHz as in the snippet above. Testing over WebRTC at 48 kHz measures a different system.
3. Add six conditions borrowed from the papers. Clean phone audio. Background noise at 10 dB SNR. Real accented speakers from your caller base. About 2% frame loss, the τ-Voice realistic rate. Overlap events: a backchannel, a real interruption and side speech, as in Full-Duplex-Bench v1.5. A 2-second tool delay, which no public benchmark tests.
4. Score end state first, then behavior. Primary metric: pass@1 against the backend state, with exact matching on spelled values, numbers and dates. Secondary metrics: yield rate on real interruptions, hold rate on backchannels and side speech, p95 response latency measured from caller-side audio, and dead air during tool calls. Report yield and hold as two numbers, never one blended score.
5. Run three trials per cell and report pass^3. That is 40 × 6 × 3 = 720 runs per configuration. Pass^3 (the share of scenarios that pass all three trials) exposes flakiness that pass@1 hides. With 40 scenarios, a single-condition pass rate near 50% carries about ±15 points of noise (1.96 × √(0.25/40)). Pool across conditions for ranking. Use per-cell results only to find failure modes.
6. Calibrate synthetic callers against humans. Synthetic callers are more patient and fluent than people, as the τ-Voice authors note. Record 20 to 30 real calls across your accent mix and confirm synthetic and real versions fail in the same places. See our synthetic callers guide and LLM-as-judge limits post.
7. Freeze the suite and re-run it on every change. Prompt edits, model upgrades, STT swaps and endpointing changes all move these numbers. Our guide to benchmarking on your own data covers versioning.
8. Get an independent read before major decisions. Before picking a model vendor or launching to a new caller population, have someone outside the build team run the suite. Evalgent does this as a pre-launch audit or vendor bake-off on your scenarios, over real phone paths, with scoring rules agreed up front.
Frequently asked questions
What is the best voice agent benchmark?
No single benchmark is best. τ-Voice is closest to a real support call because it combines verifiable tasks, domain policy, full-duplex audio and 8 kHz G.711 compression. Pair it with Full-Duplex-Bench v1.5 or HumDial for turn-taking and Audio MultiChallenge for long-call memory. Then confirm on a private benchmark built from your own calls.
What is the difference between VoiceBench and VoiceAgentBench?
VoiceBench tests whether a model follows spoken instructions as well as written ones, across knowledge, instruction following, safety and audio perturbations. It has no tools. VoiceAgentBench tests tool use from speech: tool choice, argument filling, chained calls and refusals, in English and six Indic languages. Both score text output, and both use mostly synthetic audio.
How do Full-Duplex-Bench v1, v1.5, v2 and v3 differ?
v1 scores pause handling, backchannels, turn-taking and interruptions on single exchanges. v1.5 plays overlapping speech and measures reactions and stop latency. v2 runs live multi-turn sessions with a spoken examiner and judge. v3 tests chained tool calls with real disfluent recordings. Always check which version a chart cites.
Is MultiChallenge the same as Audio MultiChallenge?
No. MultiChallenge is Scale AI's text benchmark from January 2025. OpenAI's "MultiChallenge audio" was a TTS conversion of an audio-friendly subset. Audio MultiChallenge is a separate Scale AI benchmark with 452 real recorded conversations, 1,712 rubrics and a Voice Editing axis. Scores across the three are not comparable.
Do voice agent benchmarks test phone audio?
Almost none do. τ-Voice applies G.711 μ-law at 8 kHz to every condition. The others use 16 kHz, 24 kHz or 48 kHz audio, sometimes with packet loss or low-pass filters. To see phone performance, re-run benchmark audio through an 8 kHz μ-law conversion or place real calls through your SIP trunk.
Why do cascaded pipelines look slow in benchmarks?
The cascaded baselines in papers are simple: Whisper, a text LLM and a TTS engine, often without streaming. Full-Duplex-Bench-v3's took 10.12 seconds. Production pipelines use streaming STT with end-of-turn detection and streaming TTS. Public benchmarks have not tested a tuned streaming cascade, so their latency results don't apply to one.
How many test scenarios does a private voice benchmark need?
Start with 40 scenarios across five families, run under six conditions with three trials each. That gives 720 runs per configuration. At 40 scenarios, a single-condition pass rate near 50% has about ±15 points of noise. Pool conditions for rankings, and grow toward 100 or more scenarios before you treat small gaps as real.
Can I trust vendor-reported benchmark scores?
Treat them as a shortlist signal. Check the version, the settings and the domain. GPT-Realtime-2's gains were reported at high and xhigh reasoning effort, while the default is low. Big Bench Audio changed its scoring in March 2026. On τ-Voice, one model scored 71% in Retail and 28% in Telecom.
The bottom line
Public voice agent benchmarks are good at ruling models out and poor at ruling them in, because each one measures a slice under conditions that rarely match an 8 kHz call with your entities, tools and policy. Use τ-Voice, Full-Duplex-Bench, HumDial and Audio MultiChallenge to build a shortlist, then decide with a private benchmark run over your real phone path.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more