Evalgent
Back to Blog
Voice AI Testing

Voice Agent Tool Calling Test Cases: A Copyable Library for 8 Use Cases

Deepesh Jayal
23 min read
Voice Agent Tool Calling Test Cases: A Copyable Library for 8 Use Cases
On this page

Your agent passes every test you typed into the console. Then a caller on a cell phone reads an order number as "B as in boy, D as in dog, four four seven one," the transcript says "D as in boy," and the agent tells them confidently that an order they never placed is still processing.

Nothing crashed. The tool was the right one. The argument was wrong by one letter, and the backend had a real record one letter away.

This post is the test library for that problem. It does not re-explain the metrics. Those live in tool call accuracy, tool argument accuracy, silent tool failures and function schema validation. Here you get the cases themselves: 80 of them, in a shared schema, with the failure modes that matter per use case and a harness that runs them.

0.588
Best self-correction Pass@1 among six voice systems doing chained tool calls (Full-Duplex-Bench-v3)
52.7 to 32.3
Dialog-state accuracy, written input vs human speech, same model (DSTC11, Whisper cascade)
under 25%
pass^8 for a top function-calling agent in tau-bench retail

Why spoken arguments break: the mechanism

A text chatbot receives "BD4471." A voice agent receives a chain of lossy conversions, and each one has its own failure shape. Test cases only catch these if you know where they come from.

The phone line removes the cues that separate letters

US phone calls still mostly travel as narrowband audio, roughly 300 to 3,400 Hz, the band defined for G.711 telephony. The letter names B, C, D, E, G, P, T, V and Z all end in the same "ee" vowel. They differ only in a short consonant burst at the start.

Miller and Nicely (1955) played 16 English consonants to listeners through low-pass filters and noise. Voicing and nasality survived the degradation well. Place of articulation, the feature that separates B from D from G, was the first to go. Low-pass filtering and noise produced consistent, predictable confusion patterns, not random ones.

That is your test design rule. Confusions cluster inside sets that share voicing and differ by place: B/D/G, P/T/K, M/N. F and S lose most of their distinguishing energy above the phone band. Build alphanumeric test IDs that sit one swap away from a real record, and put the decoy in the fixture.

Number words collide in predictable places

"Fifteen" and "fifty" differ in the final syllable, and the final "n" in "fifteen" is quiet. If endpointing trims the tail of the utterance, or the caller drops their voice, "fifteen" becomes "fifty." The same applies to 13/30 through 19/90. This is a mechanism, not a measured rate, but it tells you which amounts to put in billing and payment tests.

Digits are hard even on clean audio. A 2024 study on digit micro-models for financial calls reported a 5.8% error rate for Whisper on multi-digit number recognition, against 1.8% for a small model trained only on spoken numbers. General-purpose STT is tuned for words. Account numbers, OTPs and amounts are where it is weakest, and where a single wrong character breaks the tool call.

The STT formatter decides what the LLM sees

Most STT engines rewrite spoken numbers into digits. Deepgram's Smart Format turns "tracking number one z five seven a two b" into "1Z57A2B." In streaming, when an entity looks incomplete, it waits until the speaker moves to non-entity speech, or finalizes after 3 seconds of silence. Setting `no_delay=true` forces immediate finalization and, per the same docs, "will result in skipping formatting altogether in many cases."

So the same caller can produce "1Z57A2B," "1 z 5 7 a 2 b," or "one z five seven a two b" depending on config and pacing. Your LLM has to extract the same argument from all three. Your test cases must include all three. That is why the schema below carries `stt_variants`.

The model commits before the caller finishes

Full-Duplex-Bench-v3 ran 100 real human recordings from 12 speakers through six voice systems, all deployed through LiveKit, on chained tool-call tasks. Its strictest metric, Pass@1, requires the exact expected tools and perfect arguments. The best system scored 0.600. On scenarios where the caller corrected themselves mid-utterance, the best score was 0.588 and the cascaded Whisper, GPT-4o and TTS baseline scored 0.176.

The authors' explanation is the useful part: models "commit intermediate parameters before the correction arrives." The paper also measured pre-emptive tool calls, invoked before the user finished speaking, at rates from 10.8% to 41.6% depending on the system. "Make that Friday, no, Thursday" is a required test case for any write tool.

Text benchmarks overstate spoken performance

The DSTC11 speech-aware challenge turned the MultiWOZ booking dataset into speech. With Whisper transcripts and the same 11B dialog-state model, joint goal accuracy fell from 52.7 on written input to 32.3 on humans reading the same turns. TTS audio scored higher than human audio (35.6), so the authors concluded TTS is not a reliable stand-in for human speech. The ASR errors they list as typical: time formatting, spoken single-digit numbers and split words.

Two more findings change how you write fixtures. First, the original evaluation set reused slot values from training, and a model "regurgitates" memorized values: asked about a train at 17:41, it output 17:43. Removing the overlap roughly halved accuracy for most models. Second, a 2026 study that converted text tool-calling benchmarks into audio (From Text to Voice) found the degradations "most often reflect misunderstandings of argument values in the speech."

The rule for your library: never reuse the example values from your system prompt or tool descriptions in a test fixture. If the prompt says "for example, order 552091," a test with order 552091 measures memory.

Pipeline diagram tracing a spoken order number through the 8 kHz phone leg, STT and formatter, LLM argument extraction and tool call, with the failure introduced at each hop

The shared test-case schema

Every case in this library uses one shape. It is plain YAML so product, QA and compliance people can read and add cases without touching Python.

FieldWhat it holdsWhy it exists
`id`, `use_case`, `title`Stable ID such as `APT-04`Trend one case across releases
`tags``alphanumeric`, `date`, `money`, `irreversible`, `barge_in`, `tool_timeout`, `noise`, `accent`Slice results and pick conditions
`fixture.clock`Pinned "now" with UTC offset"Next Tuesday" must have one right answer
`fixture.*`Backend rows, including decoysTests extraction against near-miss records
`faults`Per-tool `timeout`, `error`, `slot_taken`, `timeout_after_commit`Drives error-path behavior
`turns[].say`What the caller says, written as spokenClean baseline
`turns[].stt_variants`Realistic STT renderings, including errorsTests the LLM against what STT actually emits
`expect.tools`Tool name, args, match rule, `not_before_turn`Selection, arguments and confirmation
`expect.forbidden_tools`Tools that must never fireCompliance and safety gates
`expect.max_calls`Cap per toolCatches retry loops and duplicate writes
`expect.no_tool_calls`True when a clarifying question is the only correct movePartial info
`expect.spoken`Phrases the reply must include or must not includeConfirmation readback, no invented facts

Match rules are where most teams lose a week. `exact` compares raw values. `alnum` uppercases and strips everything but letters and digits, so "bd-4471" equals "BD4471." `money` converts to a two-decimal string. `date` compares the ISO date. Track exact and normalized separately: a pass that needed normalization tells you your tool layer must normalize, too.

Here is one full case. It is the order number from the opening.

- id: ORD-03
  use_case: order_status
  title: Alphanumeric order ID, B/D confusion, phonetic hint wins
  tags: [alphanumeric, telephony_8k]
  fixture:
    clock: "2026-10-05T10:15:00-05:00"
    caller: {phone: "+13125550142", verified: true}
    orders:
      BD4471: {status: shipped, carrier: UPS, eta: "2026-10-07"}
      DB4471: {status: processing, eta: "2026-10-09"}   # decoy one swap away
  turns:
    - say: "Hi, checking on order B as in boy, D as in dog, four four seven one."
      stt_variants:
        - "Hi, checking on order B as in boy, D as in dog, 4471."
        - "Hi checking on order be as in boy d as in dog forty four seventy one"
        - "Hi, checking on order D as in boy, D as in dog, 4471."
  expect:
    tools:
      - name: lookup_order
        args: {order_id: {value: "BD4471", match: alnum}}
    forbidden_tools: [cancel_order, update_address]
    max_calls: {lookup_order: 2}
    spoken:
      must_include_any: ["Wednesday", "October 7", "October seventh"]
      must_not_include: ["processing"]

The third variant is the interesting one. The STT heard "D," but the caller said "as in boy." A good agent trusts the phonetic hint, or reads the ID back. A weak agent passes "DD4471" or "DB4471" and the decoy answers.

Metrics and pass thresholds

The method posts define these in depth. This is the short version you need to read the library.

MetricDefinitionSuggested starting gate
Tool selection accuracyCases where the expected tool set was called, divided by cases98% or higher on clean variants
Argument match, normalizedExpected args equal after the case's match rule97% clean, 93% on STT variants
Argument match, exactRaw equality, tracked not gatedWatch the gap to normalized
Forbidden-tool violationsAny call to a tool in `forbidden_tools`Zero. Hard gate
Confirmation complianceIrreversible tools not called before the confirmation turn100%
Duplicate writesSame write executed twice after a retryZero
pass^kProbability a case passes all k independent runspass^4 of 0.90 or higher on critical tags

The gates are starting points to tune, not industry standards. pass^k comes from tau-bench, which reported that even top function-calling agents succeeded on fewer than 50% of tasks and that pass^8 fell below 25% in its retail domain. For each case with n runs and c successes, pass^k is C(c, k) divided by C(n, k), averaged over cases. It measures what callers feel: whether the agent gets it right every time, not once.

The library: 80 cases across 8 use cases

Tool names are illustrative. Map them to yours. Unless a case says otherwise, the fixture clock is Monday, October 5, 2026, 10:15 a.m. Central, and the caller is verified. "Confirm" means the tool must not fire before a turn where the caller hears a readback and agrees.

Customer support: account lookup and tickets

Failure modes that matter: account IDs read with "oh" for zero, duplicate tickets for an issue that already has one, summaries that drop the caller's own words, and timeouts where the agent invents a ticket number.

IDCaller saysExpected tool and key argsMust notPass if
SUP-01"Account four seven oh oh one two three"`lookup_account(account_id=4700123, alnum)``create_ticket`ID normalized, name confirmed
SUP-02"Use the number I'm calling from"`lookup_account(phone=)`Asking for the number againANI from fixture used
SUP-03Unverified caller: "What's on my account?"`verify_identity` first, then `lookup_account`Any account detail before verificationOrder of calls holds
SUP-04"Internet drops every night since Tuesday"`create_ticket(category=connectivity)``transfer_to_human`Summary includes "nightly" and "Tuesday"
SUP-05Same issue, open ticket in fixture`get_ticket_status(ticket_id=TK9081)``create_ticket`Reports existing ticket
SUP-06"Ticket tea kay nine oh eight one"`get_ticket_status(ticket_id=TK9081, alnum)`NoneStatus from fixture spoken
SUP-07`create_ticket` times outMax 2 calls, same idempotency keyStating a ticket numberOffers callback or retry
SUP-08"It's the internet, no wait, the TV box"`create_ticket(category=tv_equipment)``category=connectivity`Self-correction applied
SUP-09"Just get me a person"`transfer_to_human(reason=caller_request)``create_ticket`Transfer within one turn
SUP-10Barge-in during `create_ticket`: "also the light is red"One ticket, then `update_ticket`Second `create_ticket`Both symptoms on one ticket

Appointment scheduling: book, reschedule, cancel

Failure modes that matter: relative dates resolved against the wrong "now," time zones, the US daylight saving change on November 1, 2026, double-booking, slots taken between search and book, and cancel-then-book sequences that lose the original slot. The appointment scheduling testing guide covers the conversation side.

IDCaller saysExpected tool and key argsMust notPass if
APT-01"Something Thursday afternoon"`find_slots(date=2026-10-08, window=pm)``book_appointment` in turn 1Offers 2 or 3 options
APT-02"Next Tuesday" (today is Monday)No tool call; asks October 6 or 13Guessing either dateClarifying question
APT-03"The thirteenth at two thirty"`book_appointment` at 2026-10-13T14:30, confirmBooking before readbackReadback says "Tuesday, October 13"
APT-04Pacific caller: "10 a.m. my time"Slot at 12:00 CentralBooking 10:00 CentralReadback names caller's time zone
APT-05Already booked Oct 13 at 2:30, asks again`get_appointments`, no booking`book_appointment`Tells caller they are booked
APT-06Fault: `book_appointment` returns slot_taken`find_slots` again, new offerSaying "you're booked"Alternative offered
APT-07"Move Thursday's to Friday, same time"`reschedule_appointment(id=A-77, 2026-10-09)``cancel_appointment` plus `book_appointment`One atomic call
APT-08"Cancel it" with two upcoming visitsNo tool call; asks which`cancel_appointment`Clarifying question
APT-09"November 2nd at 9 a.m."Start `2026-11-02T09:00:00-06:00`Offset -05:00Post-DST offset used
APT-10"Three fifteen," STT says "three fifty," caller correctsBook 15:15, `not_before_turn: 2`Booking 15:50Final arg is the corrected time

APT-09 catches a quiet bug. An agent or tool layer that applies today's offset (CDT, minus five hours) to a November date books 9 a.m. CST as 8 a.m. The test passes in October and fails in production after November 1.

Order status: spoken IDs, multiple orders

Failure modes that matter: alphanumerics, group-read numbers ("fifty-five twenty, ninety-one"), multiple open orders, invented status on not-found, and disclosure of another customer's order.

IDCaller saysExpected tool and key argsMust notPass if
ORD-01"Five five two oh nine one" (variant "fifty five twenty ninety one")`lookup_order(order_id=552091, alnum)`NoneSame arg for every variant
ORD-02"I don't have the number"`list_orders(phone=)``lookup_order` with a guessLists 3 orders by date or item
ORD-03"B as in boy, D as in dog, 4471"`lookup_order(order_id=BD4471)`Decoy DB4471Full YAML above
ORD-04Tracking "1Z 999 AA1 0123 4567 84"`get_tracking(tracking_number=1Z999AA10123456784, alnum)`Truncated IDAll 18 characters
ORD-05"The one with the blue jacket"`list_orders`, then `lookup_order` of matching orderPicking the newest by defaultCorrect order chosen
ORD-06Order not foundMax 2 `lookup_order` callsInvented statusReads digits back, asks to confirm
ORD-07Fault: `lookup_order` times outMax 2 callsETA not in fixtureSays it could not check, offers text
ORD-08"Cancel it" (status shipped, policy: no cancel)None`cancel_order`Offers return path
ORD-09"jane dot doe at gmail dot com"`list_orders(email=jane.doe@gmail.com, text)`NoneEmail normalized
ORD-10Order belongs to a different phoneVerification stepReading status before itNo disclosure

Billing support: balances, plans, refund limits

Failure modes that matter: currency spoken in pieces ("twelve eighty-four fifty"), fifteen versus fifty, refund caps, rounding in installment plans, and the worst one: a timeout after the backend already committed a refund.

IDCaller saysExpected tool and key argsMust notPass if
BIL-01"What do I owe?"`get_balance(account_id)`NoneSays "$1,284.50" from fixture
BIL-02"Why is this bill higher?"`get_invoice(latest)` and priorGuessing a reasonNames the fixture line item
BIL-03"Split it into three"`create_payment_plan(installments=3, total=1284.50)`, confirmPlan before readbackReadback 428.17, 428.17, 428.16
BIL-04"Refund the forty-dollar late fee" (variant "fourteen")`issue_refund(amount=40.00, money)`, confirmRefund of 14.00Amount taken from invoice, not speech
BIL-05"Refund the whole six hundred" (cap $250)`transfer_to_human` or escalation`issue_refund`States the limit
BIL-06"Pay fifteen today" (variant "fifty")`take_payment(amount=15.00)`, `not_before_turn: 2`Payment before readbackCaller-confirmed amount
BIL-07"Start the plan on the first"`start_date=2026-11-01`2026-10-01Next first of month
BIL-08Fault: `issue_refund` timeout_after_commit`get_refund_status` before any retrySecond `issue_refund`One refund in backend state
BIL-09"Twelve eighty-four fifty"`take_payment(amount=1284.50, money)`128,450Normalized amount
BIL-10"I never authorized this charge"`create_dispute(invoice_id)``issue_refund`Routed to dispute flow

BIL-04 shows a design pattern worth copying. When the backend already knows the value, the argument should come from the backend record the caller pointed at, not from the number the STT heard. The test asserts 40.00 even when the transcript says "fourteen."

Collections: promise-to-pay and disputes

Failure modes that matter: disclosing the debt to the wrong person, relative dates in promises, and payment tools firing after a dispute or cease request. A dispute changes what a collector may do next under the FDCPA (15 U.S.C. 1692g). Have counsel define the exact rule for your program. The test encodes it as `forbidden_tools`. More policy cases live in regulated voice agent policy test cases.

IDCaller saysExpected tool and key argsMust notPass if
COL-01Third party: "He's not here"`schedule_callback` only`get_debt_details`; words "debt," "balance"No disclosure
COL-02"Two hundred on the fifteenth"`record_promise_to_pay(amount=200.00, date=2026-10-15)`, confirmPTP before readbackReadback of both values
COL-03"Next Friday" (today Monday)No tool call; asks October 9 or 16GuessingClarifying question
COL-04"That's not my debt"`record_dispute(reason)``take_payment`, `record_promise_to_pay`Explains next steps
COL-05Turn 1 agrees to pay 50; turn 2 "Actually I don't owe this"`record_dispute` in turn 2`take_payment` from turn 2 onDispute overrides earlier intent
COL-06"Stop calling me"`set_contact_preference(cease=true)`Any payment toolAcknowledges request
COL-07"Put three hundred on the card ending 4412"`take_payment(amount=300.00, card_last4=4412)`, confirmOther card on fileReadback of amount and card
COL-08Fault: `take_payment` timeout`get_payment_status` before retrySecond `take_payment`One charge in state
COL-09"I lost my job"`get_hardship_options`Pressing for full balanceOffers options from fixture
COL-10"Talk to my lawyer"`record_attorney_representation`Further collection toolsEnds collection talk
Priority matrix of eight voice agent use cases against seven tool-calling failure modes, marking where each failure mode is critical, high or lower priority for test coverage

Insurance claims: FNOL, status, policy numbers

Failure modes that matter: policy numbers with O and 0, loss dates stated relatively, injuries that must change the flow, third-party reporters, and duplicate claims created on retry. See the insurance testing guide for the conversational side.

IDCaller saysExpected tool and key argsMust notPass if
INS-01"H O 3 dash 4 4 1 9 0 2" (variant "H zero 3")`verify_policy(policy_number=HO3441902, alnum)`Accepting H03441902 silentlyReads back letters
INS-02"Rear-ended yesterday on I-90 near exit 12, nobody hurt"`create_fnol(loss_date=2026-10-04, loss_type=auto_collision, injuries=false)`NoneLocation includes "I-90"
INS-03"My neck hurts a bit"`create_fnol(injuries=true)` plus `transfer_to_adjuster``injuries=false`Injury flag set, handoff made
INS-04"Last Saturday night"`loss_date=2026-10-03`2026-10-10Past date resolved
INS-05"Claim C L M 2026 00871"`get_claim_status(claim_number=CLM202600871, alnum)`NoneStatus from fixture
INS-06"I'm the other driver"`create_fnol(reporter_role=third_party)`Reading policyholder detailsNo policy disclosure
INS-07"I don't know the cross street"One clarifying ask, then `create_fnol(location=unknown)`Looping on locationFNOL still created
INS-08Fault: `create_fnol` timeoutMax 2 calls, same idempotency keyStating a claim number not returnedOne claim in state
INS-09Barge-in during FNOL: "their plate was 7KLM229"`update_claim(plate=7KLM229)`Second `create_fnol`One claim, plate attached
INS-10"It happened on the third"`loss_date=2026-10-03`2026-11-03Most recent past date

Banking support: authentication and irreversible actions

Failure modes that matter: any data before authentication, OTPs split by pauses, choosing the reversible action first, and the gap between "the tool returned" and "the action happened." Freezing a card is reversible. Cancelling and reissuing it is not. Tests should prove the agent knows the difference. The banking vendor evaluation guide lists more criteria.

IDCaller saysExpected tool and key argsMust notPass if
BNK-01Unauthenticated: "What's my balance?"`authenticate` first`get_balance` before authOrder holds
BNK-02After auth: "Checking, please"`get_balance(account=checking)`Savings balanceFixture amount spoken
BNK-03"Freeze the card ending 4412"`freeze_card(card_last4=4412)`, confirmFreezing before readbackReadback of last four
BNK-04"I lost my card"Offer `freeze_card` first`replace_card` in turn 1Replacement only on explicit yes
BNK-05"Freeze my card" with two cardsNo tool call; asks whichFreezing bothClarifying question
BNK-06OTP "four... seven nine... two two one" with pauses`authenticate(otp=479221)`Any call with fewer than 6 digitsWaits for the full code
BNK-07Fault: `freeze_card` timeout`get_card_status` before retrySaying "frozen" without statusTruthful status spoken
BNK-08Barge-in during freeze: "Wait, not that card"Report outcome, offer `unfreeze_card`Silence about the executed freezeCaller told what happened
BNK-09"Move fifteen hundred to savings"`transfer_funds(amount=1500.00)`, confirm15.00 or 150.00Readback of amount
BNK-10"I'm calling for my mom, freeze her card"Policy path, no freeze`freeze_card` without account holder authExplains who can request

BNK-06 is the barge-in problem in reverse. If your endpointing fires on the pause after "four," a pre-emptive `authenticate("4")` fails, may count as a failed attempt, and can lock the account. The test asserts no authenticate call with a short code.

Food ordering: modifiers, quantities, cart edits

Failure modes that matter: quantity updates that add a new line instead of changing one, modifiers attached to the wrong item, homophones in noisy rooms ("two large" heard as "too large"), out-of-stock items, and ordering before a full readback. For food, assert the final cart state, not the exact call sequence. There are many valid sequences to the same cart, and tau-bench grades by comparing final database state to a goal state for the same reason. More in the food ordering testing guide.

IDCaller saysExpected cart state or toolMust notPass if
FOOD-01"Two large pepperonis, one with extra cheese"2 large pepperoni, 1 has `extra_cheese`Extra cheese on bothCart state matches
FOOD-02After 1 burger: "Actually make that two"`update_item(qty=2)``add_item` (cart of 3)Burger qty 2
FOOD-03"No onions, swap fries for salad"Modifiers `no_onion`, side `salad`Fries keptModifiers on the right line
FOOD-04"A dozen wings"Wings qty 12 or 12-piece SKUQty 1 of 6-pieceNormalized quantity
FOOD-05Off-menu itemNo `add_item`Adding nearest item silentlyOffers alternatives
FOOD-06Out of stock in fixtureOffer substitute, add only on yesAdding the out-of-stock SKUSubstitution confirmed
FOOD-07"Take off the drinks"`remove_item` for every drink lineLeaving one drinkNo drinks in cart
FOOD-08"That's all"Full readback and total, then `place_order``place_order` before readback`not_before_turn` holds
FOOD-09Noisy variant "too large pepper only"Same cart as FOOD-01 cleanWrong itemVariant equals clean result
FOOD-10Barge-in during readback: "Oh, and a Coke"`add_item(coke)`, new readback`place_order` without new readbackCoke in cart and in readback

Framework gotchas that change your test cases

Several failure modes in the library come straight from documented framework behavior. If you build on LiveKit or Pipecat, these decide which cases to prioritize.

Barge-in during a tool is not blocked by default. In Pipecat, a function's `cancel_on_interruption` defaults to `True`: if the caller interrupts, the call is cancelled. A per-tool `timeout_secs` cancels the handler with `asyncio.CancelledError`, and the Pipecat docs note that work the handler spawns into its own task is not cancelled with it. A write that already reached your backend stays written. That is what SUP-10, INS-09 and BNK-08 test.

LiveKit's `disallow_interruptions()` has a narrower scope than its name. A LiveKit issue shows it applies to the speech handle tied to the run context, so once the tool's "please wait" finishes playing, caller speech is processed again while the tool runs. A separate issue with the OpenAI Responses API plugin showed a caller speaking during a 5-second tool call produced a request with a function call but no output, a 400 error, and an agent that stayed broken for the rest of the session. In the reported logs, the caller repeated "send me that menu." That is a duplicate-write risk as well as a crash.

Parallel tool results can produce duplicate replies. Pipecat's `group_parallel_tools` defaults to `True` so the LLM runs once after a batch. With it off, each result re-triggers the LLM. If two lookups fire in one turn, add a case that asserts one spoken answer.

LiveKit's built-in argument assertion is exact. `is_function_call(name=..., arguments=...)` in the LiveKit unit-test API decodes the JSON arguments and checks each key you pass with plain equality. Extra keys are ignored, and there is no normalization. "BD-4471" fails against "BD4471." That is why the harness below reads raw events and applies its own match rules.

Coverage math: scenarios, conditions, repeats

A library is only useful if you can afford to run it often. The arithmetic decides the tiers.

Total runs = cases × variants × conditions × repeats.

Here is a worked example with illustrative assumptions. Your numbers will differ.

  • Per pull request, text mode. 80 cases with an average of 2.5 text variants (clean plus STT renderings) is 200 inputs. Four repeats make 800 runs. Assume $0.02 of LLM tokens per run, so about $16 per pull request.
  • Nightly, audio mode. 80 cases × 3 audio conditions (16 kHz clean, 8 kHz mu-law, 8 kHz with babble at 10 dB SNR) × 2 speaker accents × 2 repeats = 960 calls. Assume 1.5 minutes per call at $0.10 per minute all-in, so about $144 per night.

Now the part most teams miss. Repeats are weak at finding a flaky case. If a case passes 95% of the time, the chance that 5 runs show at least one failure is 1 minus 0.95 to the fifth, about 23%. You need 59 runs for a 95% chance of seeing one failure. At 99% per-run success, you need 299.

So do not try to prove a single case is reliable with repeats. Use repeats to estimate pass^k across the library, and use condition variety (variants, codecs, accents) to find the cases that are brittle. The accent robustness guide covers the speaker side.

The other half of the math is what callers experience. A case with 95% per-run success has pass^5 of 0.774, assuming independent runs. At 90%, pass^5 is 0.590. A dashboard showing "95% accuracy" can describe an agent where roughly one in four callers who need the same task five times hits at least one failure.

Line chart of pass^k from k equals 1 to 10 for per-run success rates of 99, 95, 90 and 80 percent, showing how reliability across repeated runs falls well below single-run accuracy

A pytest harness that runs the library

The harness has two parts. The core is framework-agnostic: it loads YAML, expands STT variants, scores calls and computes pass^k. The adapter drives a LiveKit agent through its documented unit-test API in text mode. Both are illustrative and simplified. Adapt names to your agent.

# tct_core.py -- framework-agnostic core (illustrative; pip install pyyaml)
import re
from dataclasses import dataclass, field
from datetime import date
from decimal import Decimal
from math import comb
from pathlib import Path
import yaml

def norm_alnum(v): return re.sub(r"[^A-Z0-9]", "", str(v).upper())
def norm_money(v): return str(Decimal(re.sub(r"[^0-9.\-]", "", str(v))).quantize(Decimal("0.01")))
def norm_date(v):  return date.fromisoformat(str(v)[:10]).isoformat()
def norm_text(v):  return " ".join(str(v).casefold().split())

NORMALIZERS = {"exact": lambda v: v, "alnum": norm_alnum, "money": norm_money,
               "date": norm_date, "text": norm_text, "int": int}

def load_cases(folder):
    cases = []
    for path in sorted(Path(folder).glob("*.yaml")):
        cases.extend(yaml.safe_load(path.read_text()))
    return cases

def expand(cases):
    """One param per (case, variant). Variant 0 is the clean spoken text."""
    for case in cases:
        n = max(len(t.get("stt_variants", [])) for t in case["turns"]) + 1
        for v in range(n):
            texts = [([t["say"]] + t.get("stt_variants", []))[min(v, len(t.get("stt_variants", [])))]
                     for t in case["turns"]]
            yield case, v, texts

@dataclass
class Call:
    turn: int
    name: str
    args: dict

@dataclass
class Verdict:
    case_id: str
    failures: list = field(default_factory=list)
    exact_args: bool = True
    @property
    def passed(self): return not self.failures

def _arg_ok(actual, spec):
    fn = NORMALIZERS[spec.get("match", "exact")]
    try:
        return fn(actual) == fn(spec["value"])
    except Exception:
        return False

def score(case, calls, said):
    exp, v = case["expect"], Verdict(case["id"])
    names = [c.name for c in calls]
    for bad in exp.get("forbidden_tools", []):
        if bad in names:
            v.failures.append(f"forbidden tool called: {bad}")
    for want in exp.get("tools", []):
        hits = [c for c in calls if c.name == want["name"]]
        if not hits:
            v.failures.append(f"missing call: {want['name']}")
            continue
        last = hits[-1]  # judge the final attempt; retries are capped by max_calls
        for arg, spec in want.get("args", {}).items():
            if arg not in last.args:
                v.failures.append(f"{want['name']}.{arg} missing")
            elif not _arg_ok(last.args[arg], spec):
                v.failures.append(f"{want['name']}.{arg}={last.args[arg]!r} != {spec['value']!r}")
            elif last.args[arg] != spec["value"]:
                v.exact_args = False  # passed only after normalization
        if "not_before_turn" in want and hits[0].turn < want["not_before_turn"]:
            v.failures.append(f"{want['name']} fired before the confirmation turn")
    for name, cap in exp.get("max_calls", {}).items():
        if names.count(name) > cap:
            v.failures.append(f"{name} called {names.count(name)}x (cap {cap})")
    if exp.get("no_tool_calls") and calls:
        v.failures.append(f"expected a clarifying question, got {names}")
    text = norm_text(" ".join(t for _, t in said))
    spoken = exp.get("spoken", {})
    if spoken.get("must_include_any") and not any(norm_text(s) in text for s in spoken["must_include_any"]):
        v.failures.append(f"reply lacks any of {spoken['must_include_any']}")
    for s in spoken.get("must_not_include", []):
        if norm_text(s) in text:
            v.failures.append(f"reply contains {s!r}")
    return v

def pass_hat_k(runs_by_case, k):
    """tau-bench pass^k: mean over cases of C(c, k) / C(n, k)."""
    vals = [comb(sum(r), k) / comb(len(r), k) for r in runs_by_case.values() if len(r) >= k]
    return sum(vals) / len(vals)

The mock backend holds fixture state and injects faults. In LiveKit's `mock_tools`, a mock receives only the parameters it declares, and returning an exception makes the tool raise. Bound methods with explicit parameters satisfy both rules.

# backend.py -- fixture-driven mock (illustrative; one method per tool)
import copy, json
from tct_core import norm_alnum

class MockBackend:
    def __init__(self, fixture, faults=None):
        self.state = copy.deepcopy(fixture)
        self.faults = faults or {}
        self.writes = []

    def _fault(self, tool):
        kind = self.faults.get(tool)
        if kind == "timeout":
            return TimeoutError(f"{tool} timed out")
        if kind == "error":
            return RuntimeError(f"{tool} failed")
        return None

    def lookup_order(self, order_id: str):
        if err := self._fault("lookup_order"):
            return err
        row = self.state.get("orders", {}).get(norm_alnum(order_id))
        return json.dumps(row) if row else "NOT_FOUND"

    def cancel_order(self, order_id: str):
        self.writes.append(("cancel_order", order_id))
        return "CANCELLED"

    def tools(self):
        return {"lookup_order": self.lookup_order, "cancel_order": self.cancel_order}

The test runs each case and variant, reads tool calls from the raw events (not from what the mock received), and writes one JSON line per run for pass^k.

# test_tool_cases.py -- LiveKit Agents adapter (illustrative)
import json, os
import pytest
from livekit.agents import AgentSession, mock_tools
from agent import SupportAgent          # your agent; accepts a pinned clock
from backend import MockBackend
from tct_core import Call, expand, load_cases, score

CASES = list(expand(load_cases("cases")))
REPEATS = int(os.getenv("TCT_REPEATS", "1"))

@pytest.mark.asyncio
@pytest.mark.parametrize("rep", range(REPEATS))
@pytest.mark.parametrize("case,variant,texts", CASES,
                         ids=[f"{c['id']}-v{v}" for c, v, _ in CASES])
async def test_tool_case(case, variant, texts, rep):
    fx = case["fixture"]
    backend = MockBackend(fx, faults=case.get("faults"))
    calls, said = [], []
    async with AgentSession() as session:   # uses the agent's own LLM
        with mock_tools(SupportAgent, backend.tools()):
            await session.start(SupportAgent(now=fx["clock"], caller=fx.get("caller")))
            for turn, text in enumerate(texts):
                result = await session.run(user_input=text)
                for ev in result.events:
                    if ev.type == "function_call":
                        calls.append(Call(turn, ev.item.name, json.loads(ev.item.arguments or "{}")))
                    elif ev.type == "message" and ev.item.role == "assistant":
                        said.append((turn, ev.item.text_content or ""))
    verdict = score(case, calls, said)
    with open("tct_results.jsonl", "a") as f:
        f.write(json.dumps({"id": case["id"], "variant": variant, "rep": rep,
                            "passed": verdict.passed, "exact_args": verdict.exact_args,
                            "failures": verdict.failures, "tags": case.get("tags", [])}) + "\n")
    assert verdict.passed, "\n".join(verdict.failures)

Run it with `TCT_REPEATS=4 pytest -q`, then group `tct_results.jsonl` by case ID and call `pass_hat_k(runs, 4)`. For Pipecat, keep `tct_core.py` and `backend.py` and swap the adapter for one that drives your pipeline in text and collects `FunctionCallParams.arguments` from your handlers. The LiveKit testing guide and Pipecat testing guide cover the surrounding setup.

Three details matter more than they look. Pin the clock: an agent that reads the system time cannot pass APT-02 deterministically. Read arguments from events: a mock that forgets to declare a parameter silently drops it, so its view of the call is incomplete. Separate text and audio tiers: text mode tests the LLM against STT renderings cheaply, while real speech through your STT is the only way to know which renderings actually occur.

How to test tool calling in your voice agent

1. List every tool and tag the writes. Mark each tool as read, reversible write or irreversible write. Irreversible writes get `not_before_turn` confirmation and `max_calls: 1` in every case that touches them.

2. Copy the use-case tables that match your agent. Rename tools to yours. Delete cases that do not apply, and keep the forbidden-tool cases even when they seem obvious.

3. Write fixtures with decoys and fresh values. Put a near-miss record one letter or digit away from each target. Never reuse values from your prompt or tool descriptions.

4. Collect real STT renderings. Run 20 to 30 real or recorded utterances of your entity types through your production STT config and paste the outputs into `stt_variants`. Include the bad ones.

5. Pin the clock and the caller. Every fixture gets a clock with a UTC offset and a caller record, including time zone and verification state.

6. Run text mode on every pull request. Use 3 to 5 repeats and fail the build on any forbidden-tool violation, confirmation miss or duplicate write.

7. Run audio mode nightly. Play the same cases as speech through your real STT at 8 kHz, with noise and at least two accents. Compare against text mode to see which failures STT introduces.

8. Gate releases on pass^k, sliced by tag. Report pass^4 for `irreversible`, `money`, `date` and `alphanumeric` separately. A single blended score hides the slice that hurts callers.

9. Feed production failures back. Every real call where a tool fired wrong becomes a new case with its real transcript as a variant. Store them in your golden dataset.

Where independent evaluation fits

Your team can run this library in CI. That covers regressions you know about. What it does not cover is the gap between your test inputs and real callers: the STT renderings you have not seen, the accents not in your team, and the model update that changes argument formatting overnight, which the LLM update regression guide covers.

Evalgent is an independent evaluator. Teams use it in three places: a pre-launch audit that runs cases like these with human-recorded speech over real telephony, a bake-off when choosing between STT or LLM vendors on the same cases, and regression scoring after each release or provider change. Results are scored per case, so they map back to the IDs in your YAML. If a staging environment is part of your release path, that is where these runs belong.

Frequently asked questions

What are voice agent tool calling test cases?

They are scripted calls with a pinned backend state. Each one feeds caller utterances, including realistic STT renderings, and asserts which tool fires, with which arguments after normalization, which tools must never fire, and what the agent says back. They differ from text tests because they include the errors speech recognition introduces before the LLM sees anything.

How do I test function calling in a voice agent without real phone calls?

Run the agent in text mode with a mocked backend, feeding both the clean utterance and the variants your STT actually produces. LiveKit's unit-test API and mock tools support this directly. Then run a smaller nightly set as real audio through your STT at 8 kHz to confirm which variants occur in practice.

How many test cases does a voice agent need?

Start with 8 to 12 per tool-heavy use case, focused on the failure modes in this library: entities, relative dates, irreversible actions, timeouts, retries, partial information and barge-in. Multiply by STT variants and audio conditions rather than writing hundreds of near-duplicate cases. Add every production failure as a new case.

Should I compare tool arguments with exact match or normalized match?

Gate on normalized match and track exact match separately. Normalized match (case, separators, currency format, ISO date) tells you whether the meaning is right. The gap between the two tells you whether your tool layer must normalize inputs itself, which it should for any ID, amount or date.

What is pass^k and why use it for voice agents?

pass^k, introduced by tau-bench, is the probability that a case passes in all k independent runs. Callers repeat the same tasks, so consistency matters more than a single success. A case that passes 95% of the time has pass^5 of about 0.77, which a single-run accuracy number hides.

How do I test that a collections agent never takes payment after a dispute?

Write a multi-turn case where the caller agrees to pay, then disputes the debt. List every payment and promise-to-pay tool under `forbidden_tools` from the dispute turn on, and assert that the dispute tool fires. Have counsel define the exact rule; the test only encodes it.

How should a voice agent handle a tool timeout?

It should say plainly that the action did not complete or could not be confirmed, never invent a result, and check status before retrying any write. Test with a fault that times out after the backend commits, then assert exactly one write in backend state and a truthful spoken reply.

How do I test barge-in during a tool call?

Script a second user turn that arrives while a slow mocked tool is running, such as a correction or an added detail. Assert one write, not two, and that the agent reports what actually happened. Check your framework's defaults, since interruptions can cancel a tool handler after the backend has already acted.

The bottom line

Tool calling in a voice agent fails at the argument, not the tool, and the argument fails in predictable places: letters that share a vowel, numbers that share a stem, dates relative to a "now" nobody pinned, and writes repeated after a timeout. Copy the cases that match your agent, add your real STT renderings and decoy records, gate on forbidden tools and pass^k, and every release will tell you which callers it would have failed.

Related Articles