Voice Agent Tool Calling Test Cases: A Copyable Library for 8 Use Cases

On this page
Your agent passes every test you typed into the console. Then a caller on a cell phone reads an order number as "B as in boy, D as in dog, four four seven one," the transcript says "D as in boy," and the agent tells them confidently that an order they never placed is still processing.
Nothing crashed. The tool was the right one. The argument was wrong by one letter, and the backend had a real record one letter away.
This post is the test library for that problem. It does not re-explain the metrics. Those live in tool call accuracy, tool argument accuracy, silent tool failures and function schema validation. Here you get the cases themselves: 80 of them, in a shared schema, with the failure modes that matter per use case and a harness that runs them.
Why spoken arguments break: the mechanism
A text chatbot receives "BD4471." A voice agent receives a chain of lossy conversions, and each one has its own failure shape. Test cases only catch these if you know where they come from.
The phone line removes the cues that separate letters
US phone calls still mostly travel as narrowband audio, roughly 300 to 3,400 Hz, the band defined for G.711 telephony. The letter names B, C, D, E, G, P, T, V and Z all end in the same "ee" vowel. They differ only in a short consonant burst at the start.
Miller and Nicely (1955) played 16 English consonants to listeners through low-pass filters and noise. Voicing and nasality survived the degradation well. Place of articulation, the feature that separates B from D from G, was the first to go. Low-pass filtering and noise produced consistent, predictable confusion patterns, not random ones.
That is your test design rule. Confusions cluster inside sets that share voicing and differ by place: B/D/G, P/T/K, M/N. F and S lose most of their distinguishing energy above the phone band. Build alphanumeric test IDs that sit one swap away from a real record, and put the decoy in the fixture.
Number words collide in predictable places
"Fifteen" and "fifty" differ in the final syllable, and the final "n" in "fifteen" is quiet. If endpointing trims the tail of the utterance, or the caller drops their voice, "fifteen" becomes "fifty." The same applies to 13/30 through 19/90. This is a mechanism, not a measured rate, but it tells you which amounts to put in billing and payment tests.
Digits are hard even on clean audio. A 2024 study on digit micro-models for financial calls reported a 5.8% error rate for Whisper on multi-digit number recognition, against 1.8% for a small model trained only on spoken numbers. General-purpose STT is tuned for words. Account numbers, OTPs and amounts are where it is weakest, and where a single wrong character breaks the tool call.
The STT formatter decides what the LLM sees
Most STT engines rewrite spoken numbers into digits. Deepgram's Smart Format turns "tracking number one z five seven a two b" into "1Z57A2B." In streaming, when an entity looks incomplete, it waits until the speaker moves to non-entity speech, or finalizes after 3 seconds of silence. Setting `no_delay=true` forces immediate finalization and, per the same docs, "will result in skipping formatting altogether in many cases."
So the same caller can produce "1Z57A2B," "1 z 5 7 a 2 b," or "one z five seven a two b" depending on config and pacing. Your LLM has to extract the same argument from all three. Your test cases must include all three. That is why the schema below carries `stt_variants`.
The model commits before the caller finishes
Full-Duplex-Bench-v3 ran 100 real human recordings from 12 speakers through six voice systems, all deployed through LiveKit, on chained tool-call tasks. Its strictest metric, Pass@1, requires the exact expected tools and perfect arguments. The best system scored 0.600. On scenarios where the caller corrected themselves mid-utterance, the best score was 0.588 and the cascaded Whisper, GPT-4o and TTS baseline scored 0.176.
The authors' explanation is the useful part: models "commit intermediate parameters before the correction arrives." The paper also measured pre-emptive tool calls, invoked before the user finished speaking, at rates from 10.8% to 41.6% depending on the system. "Make that Friday, no, Thursday" is a required test case for any write tool.
Text benchmarks overstate spoken performance
The DSTC11 speech-aware challenge turned the MultiWOZ booking dataset into speech. With Whisper transcripts and the same 11B dialog-state model, joint goal accuracy fell from 52.7 on written input to 32.3 on humans reading the same turns. TTS audio scored higher than human audio (35.6), so the authors concluded TTS is not a reliable stand-in for human speech. The ASR errors they list as typical: time formatting, spoken single-digit numbers and split words.
Two more findings change how you write fixtures. First, the original evaluation set reused slot values from training, and a model "regurgitates" memorized values: asked about a train at 17:41, it output 17:43. Removing the overlap roughly halved accuracy for most models. Second, a 2026 study that converted text tool-calling benchmarks into audio (From Text to Voice) found the degradations "most often reflect misunderstandings of argument values in the speech."
The rule for your library: never reuse the example values from your system prompt or tool descriptions in a test fixture. If the prompt says "for example, order 552091," a test with order 552091 measures memory.

The shared test-case schema
Every case in this library uses one shape. It is plain YAML so product, QA and compliance people can read and add cases without touching Python.
| Field | What it holds | Why it exists |
|---|---|---|
| `id`, `use_case`, `title` | Stable ID such as `APT-04` | Trend one case across releases |
| `tags` | `alphanumeric`, `date`, `money`, `irreversible`, `barge_in`, `tool_timeout`, `noise`, `accent` | Slice results and pick conditions |
| `fixture.clock` | Pinned "now" with UTC offset | "Next Tuesday" must have one right answer |
| `fixture.*` | Backend rows, including decoys | Tests extraction against near-miss records |
| `faults` | Per-tool `timeout`, `error`, `slot_taken`, `timeout_after_commit` | Drives error-path behavior |
| `turns[].say` | What the caller says, written as spoken | Clean baseline |
| `turns[].stt_variants` | Realistic STT renderings, including errors | Tests the LLM against what STT actually emits |
| `expect.tools` | Tool name, args, match rule, `not_before_turn` | Selection, arguments and confirmation |
| `expect.forbidden_tools` | Tools that must never fire | Compliance and safety gates |
| `expect.max_calls` | Cap per tool | Catches retry loops and duplicate writes |
| `expect.no_tool_calls` | True when a clarifying question is the only correct move | Partial info |
| `expect.spoken` | Phrases the reply must include or must not include | Confirmation readback, no invented facts |
Match rules are where most teams lose a week. `exact` compares raw values. `alnum` uppercases and strips everything but letters and digits, so "bd-4471" equals "BD4471." `money` converts to a two-decimal string. `date` compares the ISO date. Track exact and normalized separately: a pass that needed normalization tells you your tool layer must normalize, too.
Here is one full case. It is the order number from the opening.
- id: ORD-03
use_case: order_status
title: Alphanumeric order ID, B/D confusion, phonetic hint wins
tags: [alphanumeric, telephony_8k]
fixture:
clock: "2026-10-05T10:15:00-05:00"
caller: {phone: "+13125550142", verified: true}
orders:
BD4471: {status: shipped, carrier: UPS, eta: "2026-10-07"}
DB4471: {status: processing, eta: "2026-10-09"} # decoy one swap away
turns:
- say: "Hi, checking on order B as in boy, D as in dog, four four seven one."
stt_variants:
- "Hi, checking on order B as in boy, D as in dog, 4471."
- "Hi checking on order be as in boy d as in dog forty four seventy one"
- "Hi, checking on order D as in boy, D as in dog, 4471."
expect:
tools:
- name: lookup_order
args: {order_id: {value: "BD4471", match: alnum}}
forbidden_tools: [cancel_order, update_address]
max_calls: {lookup_order: 2}
spoken:
must_include_any: ["Wednesday", "October 7", "October seventh"]
must_not_include: ["processing"]The third variant is the interesting one. The STT heard "D," but the caller said "as in boy." A good agent trusts the phonetic hint, or reads the ID back. A weak agent passes "DD4471" or "DB4471" and the decoy answers.
Metrics and pass thresholds
The method posts define these in depth. This is the short version you need to read the library.
| Metric | Definition | Suggested starting gate |
|---|---|---|
| Tool selection accuracy | Cases where the expected tool set was called, divided by cases | 98% or higher on clean variants |
| Argument match, normalized | Expected args equal after the case's match rule | 97% clean, 93% on STT variants |
| Argument match, exact | Raw equality, tracked not gated | Watch the gap to normalized |
| Forbidden-tool violations | Any call to a tool in `forbidden_tools` | Zero. Hard gate |
| Confirmation compliance | Irreversible tools not called before the confirmation turn | 100% |
| Duplicate writes | Same write executed twice after a retry | Zero |
| pass^k | Probability a case passes all k independent runs | pass^4 of 0.90 or higher on critical tags |
The gates are starting points to tune, not industry standards. pass^k comes from tau-bench, which reported that even top function-calling agents succeeded on fewer than 50% of tasks and that pass^8 fell below 25% in its retail domain. For each case with n runs and c successes, pass^k is C(c, k) divided by C(n, k), averaged over cases. It measures what callers feel: whether the agent gets it right every time, not once.
The library: 80 cases across 8 use cases
Tool names are illustrative. Map them to yours. Unless a case says otherwise, the fixture clock is Monday, October 5, 2026, 10:15 a.m. Central, and the caller is verified. "Confirm" means the tool must not fire before a turn where the caller hears a readback and agrees.
Customer support: account lookup and tickets
Failure modes that matter: account IDs read with "oh" for zero, duplicate tickets for an issue that already has one, summaries that drop the caller's own words, and timeouts where the agent invents a ticket number.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| SUP-01 | "Account four seven oh oh one two three" | `lookup_account(account_id=4700123, alnum)` | `create_ticket` | ID normalized, name confirmed |
| SUP-02 | "Use the number I'm calling from" | `lookup_account(phone= | Asking for the number again | ANI from fixture used |
| SUP-03 | Unverified caller: "What's on my account?" | `verify_identity` first, then `lookup_account` | Any account detail before verification | Order of calls holds |
| SUP-04 | "Internet drops every night since Tuesday" | `create_ticket(category=connectivity)` | `transfer_to_human` | Summary includes "nightly" and "Tuesday" |
| SUP-05 | Same issue, open ticket in fixture | `get_ticket_status(ticket_id=TK9081)` | `create_ticket` | Reports existing ticket |
| SUP-06 | "Ticket tea kay nine oh eight one" | `get_ticket_status(ticket_id=TK9081, alnum)` | None | Status from fixture spoken |
| SUP-07 | `create_ticket` times out | Max 2 calls, same idempotency key | Stating a ticket number | Offers callback or retry |
| SUP-08 | "It's the internet, no wait, the TV box" | `create_ticket(category=tv_equipment)` | `category=connectivity` | Self-correction applied |
| SUP-09 | "Just get me a person" | `transfer_to_human(reason=caller_request)` | `create_ticket` | Transfer within one turn |
| SUP-10 | Barge-in during `create_ticket`: "also the light is red" | One ticket, then `update_ticket` | Second `create_ticket` | Both symptoms on one ticket |
Appointment scheduling: book, reschedule, cancel
Failure modes that matter: relative dates resolved against the wrong "now," time zones, the US daylight saving change on November 1, 2026, double-booking, slots taken between search and book, and cancel-then-book sequences that lose the original slot. The appointment scheduling testing guide covers the conversation side.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| APT-01 | "Something Thursday afternoon" | `find_slots(date=2026-10-08, window=pm)` | `book_appointment` in turn 1 | Offers 2 or 3 options |
| APT-02 | "Next Tuesday" (today is Monday) | No tool call; asks October 6 or 13 | Guessing either date | Clarifying question |
| APT-03 | "The thirteenth at two thirty" | `book_appointment` at 2026-10-13T14:30, confirm | Booking before readback | Readback says "Tuesday, October 13" |
| APT-04 | Pacific caller: "10 a.m. my time" | Slot at 12:00 Central | Booking 10:00 Central | Readback names caller's time zone |
| APT-05 | Already booked Oct 13 at 2:30, asks again | `get_appointments`, no booking | `book_appointment` | Tells caller they are booked |
| APT-06 | Fault: `book_appointment` returns slot_taken | `find_slots` again, new offer | Saying "you're booked" | Alternative offered |
| APT-07 | "Move Thursday's to Friday, same time" | `reschedule_appointment(id=A-77, 2026-10-09)` | `cancel_appointment` plus `book_appointment` | One atomic call |
| APT-08 | "Cancel it" with two upcoming visits | No tool call; asks which | `cancel_appointment` | Clarifying question |
| APT-09 | "November 2nd at 9 a.m." | Start `2026-11-02T09:00:00-06:00` | Offset -05:00 | Post-DST offset used |
| APT-10 | "Three fifteen," STT says "three fifty," caller corrects | Book 15:15, `not_before_turn: 2` | Booking 15:50 | Final arg is the corrected time |
APT-09 catches a quiet bug. An agent or tool layer that applies today's offset (CDT, minus five hours) to a November date books 9 a.m. CST as 8 a.m. The test passes in October and fails in production after November 1.
Order status: spoken IDs, multiple orders
Failure modes that matter: alphanumerics, group-read numbers ("fifty-five twenty, ninety-one"), multiple open orders, invented status on not-found, and disclosure of another customer's order.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| ORD-01 | "Five five two oh nine one" (variant "fifty five twenty ninety one") | `lookup_order(order_id=552091, alnum)` | None | Same arg for every variant |
| ORD-02 | "I don't have the number" | `list_orders(phone= | `lookup_order` with a guess | Lists 3 orders by date or item |
| ORD-03 | "B as in boy, D as in dog, 4471" | `lookup_order(order_id=BD4471)` | Decoy DB4471 | Full YAML above |
| ORD-04 | Tracking "1Z 999 AA1 0123 4567 84" | `get_tracking(tracking_number=1Z999AA10123456784, alnum)` | Truncated ID | All 18 characters |
| ORD-05 | "The one with the blue jacket" | `list_orders`, then `lookup_order` of matching order | Picking the newest by default | Correct order chosen |
| ORD-06 | Order not found | Max 2 `lookup_order` calls | Invented status | Reads digits back, asks to confirm |
| ORD-07 | Fault: `lookup_order` times out | Max 2 calls | ETA not in fixture | Says it could not check, offers text |
| ORD-08 | "Cancel it" (status shipped, policy: no cancel) | None | `cancel_order` | Offers return path |
| ORD-09 | "jane dot doe at gmail dot com" | `list_orders(email=jane.doe@gmail.com, text)` | None | Email normalized |
| ORD-10 | Order belongs to a different phone | Verification step | Reading status before it | No disclosure |
Billing support: balances, plans, refund limits
Failure modes that matter: currency spoken in pieces ("twelve eighty-four fifty"), fifteen versus fifty, refund caps, rounding in installment plans, and the worst one: a timeout after the backend already committed a refund.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| BIL-01 | "What do I owe?" | `get_balance(account_id)` | None | Says "$1,284.50" from fixture |
| BIL-02 | "Why is this bill higher?" | `get_invoice(latest)` and prior | Guessing a reason | Names the fixture line item |
| BIL-03 | "Split it into three" | `create_payment_plan(installments=3, total=1284.50)`, confirm | Plan before readback | Readback 428.17, 428.17, 428.16 |
| BIL-04 | "Refund the forty-dollar late fee" (variant "fourteen") | `issue_refund(amount=40.00, money)`, confirm | Refund of 14.00 | Amount taken from invoice, not speech |
| BIL-05 | "Refund the whole six hundred" (cap $250) | `transfer_to_human` or escalation | `issue_refund` | States the limit |
| BIL-06 | "Pay fifteen today" (variant "fifty") | `take_payment(amount=15.00)`, `not_before_turn: 2` | Payment before readback | Caller-confirmed amount |
| BIL-07 | "Start the plan on the first" | `start_date=2026-11-01` | 2026-10-01 | Next first of month |
| BIL-08 | Fault: `issue_refund` timeout_after_commit | `get_refund_status` before any retry | Second `issue_refund` | One refund in backend state |
| BIL-09 | "Twelve eighty-four fifty" | `take_payment(amount=1284.50, money)` | 128,450 | Normalized amount |
| BIL-10 | "I never authorized this charge" | `create_dispute(invoice_id)` | `issue_refund` | Routed to dispute flow |
BIL-04 shows a design pattern worth copying. When the backend already knows the value, the argument should come from the backend record the caller pointed at, not from the number the STT heard. The test asserts 40.00 even when the transcript says "fourteen."
Collections: promise-to-pay and disputes
Failure modes that matter: disclosing the debt to the wrong person, relative dates in promises, and payment tools firing after a dispute or cease request. A dispute changes what a collector may do next under the FDCPA (15 U.S.C. 1692g). Have counsel define the exact rule for your program. The test encodes it as `forbidden_tools`. More policy cases live in regulated voice agent policy test cases.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| COL-01 | Third party: "He's not here" | `schedule_callback` only | `get_debt_details`; words "debt," "balance" | No disclosure |
| COL-02 | "Two hundred on the fifteenth" | `record_promise_to_pay(amount=200.00, date=2026-10-15)`, confirm | PTP before readback | Readback of both values |
| COL-03 | "Next Friday" (today Monday) | No tool call; asks October 9 or 16 | Guessing | Clarifying question |
| COL-04 | "That's not my debt" | `record_dispute(reason)` | `take_payment`, `record_promise_to_pay` | Explains next steps |
| COL-05 | Turn 1 agrees to pay 50; turn 2 "Actually I don't owe this" | `record_dispute` in turn 2 | `take_payment` from turn 2 on | Dispute overrides earlier intent |
| COL-06 | "Stop calling me" | `set_contact_preference(cease=true)` | Any payment tool | Acknowledges request |
| COL-07 | "Put three hundred on the card ending 4412" | `take_payment(amount=300.00, card_last4=4412)`, confirm | Other card on file | Readback of amount and card |
| COL-08 | Fault: `take_payment` timeout | `get_payment_status` before retry | Second `take_payment` | One charge in state |
| COL-09 | "I lost my job" | `get_hardship_options` | Pressing for full balance | Offers options from fixture |
| COL-10 | "Talk to my lawyer" | `record_attorney_representation` | Further collection tools | Ends collection talk |

Insurance claims: FNOL, status, policy numbers
Failure modes that matter: policy numbers with O and 0, loss dates stated relatively, injuries that must change the flow, third-party reporters, and duplicate claims created on retry. See the insurance testing guide for the conversational side.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| INS-01 | "H O 3 dash 4 4 1 9 0 2" (variant "H zero 3") | `verify_policy(policy_number=HO3441902, alnum)` | Accepting H03441902 silently | Reads back letters |
| INS-02 | "Rear-ended yesterday on I-90 near exit 12, nobody hurt" | `create_fnol(loss_date=2026-10-04, loss_type=auto_collision, injuries=false)` | None | Location includes "I-90" |
| INS-03 | "My neck hurts a bit" | `create_fnol(injuries=true)` plus `transfer_to_adjuster` | `injuries=false` | Injury flag set, handoff made |
| INS-04 | "Last Saturday night" | `loss_date=2026-10-03` | 2026-10-10 | Past date resolved |
| INS-05 | "Claim C L M 2026 00871" | `get_claim_status(claim_number=CLM202600871, alnum)` | None | Status from fixture |
| INS-06 | "I'm the other driver" | `create_fnol(reporter_role=third_party)` | Reading policyholder details | No policy disclosure |
| INS-07 | "I don't know the cross street" | One clarifying ask, then `create_fnol(location=unknown)` | Looping on location | FNOL still created |
| INS-08 | Fault: `create_fnol` timeout | Max 2 calls, same idempotency key | Stating a claim number not returned | One claim in state |
| INS-09 | Barge-in during FNOL: "their plate was 7KLM229" | `update_claim(plate=7KLM229)` | Second `create_fnol` | One claim, plate attached |
| INS-10 | "It happened on the third" | `loss_date=2026-10-03` | 2026-11-03 | Most recent past date |
Banking support: authentication and irreversible actions
Failure modes that matter: any data before authentication, OTPs split by pauses, choosing the reversible action first, and the gap between "the tool returned" and "the action happened." Freezing a card is reversible. Cancelling and reissuing it is not. Tests should prove the agent knows the difference. The banking vendor evaluation guide lists more criteria.
| ID | Caller says | Expected tool and key args | Must not | Pass if |
|---|---|---|---|---|
| BNK-01 | Unauthenticated: "What's my balance?" | `authenticate` first | `get_balance` before auth | Order holds |
| BNK-02 | After auth: "Checking, please" | `get_balance(account=checking)` | Savings balance | Fixture amount spoken |
| BNK-03 | "Freeze the card ending 4412" | `freeze_card(card_last4=4412)`, confirm | Freezing before readback | Readback of last four |
| BNK-04 | "I lost my card" | Offer `freeze_card` first | `replace_card` in turn 1 | Replacement only on explicit yes |
| BNK-05 | "Freeze my card" with two cards | No tool call; asks which | Freezing both | Clarifying question |
| BNK-06 | OTP "four... seven nine... two two one" with pauses | `authenticate(otp=479221)` | Any call with fewer than 6 digits | Waits for the full code |
| BNK-07 | Fault: `freeze_card` timeout | `get_card_status` before retry | Saying "frozen" without status | Truthful status spoken |
| BNK-08 | Barge-in during freeze: "Wait, not that card" | Report outcome, offer `unfreeze_card` | Silence about the executed freeze | Caller told what happened |
| BNK-09 | "Move fifteen hundred to savings" | `transfer_funds(amount=1500.00)`, confirm | 15.00 or 150.00 | Readback of amount |
| BNK-10 | "I'm calling for my mom, freeze her card" | Policy path, no freeze | `freeze_card` without account holder auth | Explains who can request |
BNK-06 is the barge-in problem in reverse. If your endpointing fires on the pause after "four," a pre-emptive `authenticate("4")` fails, may count as a failed attempt, and can lock the account. The test asserts no authenticate call with a short code.
Food ordering: modifiers, quantities, cart edits
Failure modes that matter: quantity updates that add a new line instead of changing one, modifiers attached to the wrong item, homophones in noisy rooms ("two large" heard as "too large"), out-of-stock items, and ordering before a full readback. For food, assert the final cart state, not the exact call sequence. There are many valid sequences to the same cart, and tau-bench grades by comparing final database state to a goal state for the same reason. More in the food ordering testing guide.
| ID | Caller says | Expected cart state or tool | Must not | Pass if |
|---|---|---|---|---|
| FOOD-01 | "Two large pepperonis, one with extra cheese" | 2 large pepperoni, 1 has `extra_cheese` | Extra cheese on both | Cart state matches |
| FOOD-02 | After 1 burger: "Actually make that two" | `update_item(qty=2)` | `add_item` (cart of 3) | Burger qty 2 |
| FOOD-03 | "No onions, swap fries for salad" | Modifiers `no_onion`, side `salad` | Fries kept | Modifiers on the right line |
| FOOD-04 | "A dozen wings" | Wings qty 12 or 12-piece SKU | Qty 1 of 6-piece | Normalized quantity |
| FOOD-05 | Off-menu item | No `add_item` | Adding nearest item silently | Offers alternatives |
| FOOD-06 | Out of stock in fixture | Offer substitute, add only on yes | Adding the out-of-stock SKU | Substitution confirmed |
| FOOD-07 | "Take off the drinks" | `remove_item` for every drink line | Leaving one drink | No drinks in cart |
| FOOD-08 | "That's all" | Full readback and total, then `place_order` | `place_order` before readback | `not_before_turn` holds |
| FOOD-09 | Noisy variant "too large pepper only" | Same cart as FOOD-01 clean | Wrong item | Variant equals clean result |
| FOOD-10 | Barge-in during readback: "Oh, and a Coke" | `add_item(coke)`, new readback | `place_order` without new readback | Coke in cart and in readback |
Framework gotchas that change your test cases
Several failure modes in the library come straight from documented framework behavior. If you build on LiveKit or Pipecat, these decide which cases to prioritize.
Barge-in during a tool is not blocked by default. In Pipecat, a function's `cancel_on_interruption` defaults to `True`: if the caller interrupts, the call is cancelled. A per-tool `timeout_secs` cancels the handler with `asyncio.CancelledError`, and the Pipecat docs note that work the handler spawns into its own task is not cancelled with it. A write that already reached your backend stays written. That is what SUP-10, INS-09 and BNK-08 test.
LiveKit's `disallow_interruptions()` has a narrower scope than its name. A LiveKit issue shows it applies to the speech handle tied to the run context, so once the tool's "please wait" finishes playing, caller speech is processed again while the tool runs. A separate issue with the OpenAI Responses API plugin showed a caller speaking during a 5-second tool call produced a request with a function call but no output, a 400 error, and an agent that stayed broken for the rest of the session. In the reported logs, the caller repeated "send me that menu." That is a duplicate-write risk as well as a crash.
Parallel tool results can produce duplicate replies. Pipecat's `group_parallel_tools` defaults to `True` so the LLM runs once after a batch. With it off, each result re-triggers the LLM. If two lookups fire in one turn, add a case that asserts one spoken answer.
LiveKit's built-in argument assertion is exact. `is_function_call(name=..., arguments=...)` in the LiveKit unit-test API decodes the JSON arguments and checks each key you pass with plain equality. Extra keys are ignored, and there is no normalization. "BD-4471" fails against "BD4471." That is why the harness below reads raw events and applies its own match rules.
Coverage math: scenarios, conditions, repeats
A library is only useful if you can afford to run it often. The arithmetic decides the tiers.
Total runs = cases × variants × conditions × repeats.
Here is a worked example with illustrative assumptions. Your numbers will differ.
- Per pull request, text mode. 80 cases with an average of 2.5 text variants (clean plus STT renderings) is 200 inputs. Four repeats make 800 runs. Assume $0.02 of LLM tokens per run, so about $16 per pull request.
- Nightly, audio mode. 80 cases × 3 audio conditions (16 kHz clean, 8 kHz mu-law, 8 kHz with babble at 10 dB SNR) × 2 speaker accents × 2 repeats = 960 calls. Assume 1.5 minutes per call at $0.10 per minute all-in, so about $144 per night.
Now the part most teams miss. Repeats are weak at finding a flaky case. If a case passes 95% of the time, the chance that 5 runs show at least one failure is 1 minus 0.95 to the fifth, about 23%. You need 59 runs for a 95% chance of seeing one failure. At 99% per-run success, you need 299.
So do not try to prove a single case is reliable with repeats. Use repeats to estimate pass^k across the library, and use condition variety (variants, codecs, accents) to find the cases that are brittle. The accent robustness guide covers the speaker side.
The other half of the math is what callers experience. A case with 95% per-run success has pass^5 of 0.774, assuming independent runs. At 90%, pass^5 is 0.590. A dashboard showing "95% accuracy" can describe an agent where roughly one in four callers who need the same task five times hits at least one failure.

A pytest harness that runs the library
The harness has two parts. The core is framework-agnostic: it loads YAML, expands STT variants, scores calls and computes pass^k. The adapter drives a LiveKit agent through its documented unit-test API in text mode. Both are illustrative and simplified. Adapt names to your agent.
# tct_core.py -- framework-agnostic core (illustrative; pip install pyyaml)
import re
from dataclasses import dataclass, field
from datetime import date
from decimal import Decimal
from math import comb
from pathlib import Path
import yaml
def norm_alnum(v): return re.sub(r"[^A-Z0-9]", "", str(v).upper())
def norm_money(v): return str(Decimal(re.sub(r"[^0-9.\-]", "", str(v))).quantize(Decimal("0.01")))
def norm_date(v): return date.fromisoformat(str(v)[:10]).isoformat()
def norm_text(v): return " ".join(str(v).casefold().split())
NORMALIZERS = {"exact": lambda v: v, "alnum": norm_alnum, "money": norm_money,
"date": norm_date, "text": norm_text, "int": int}
def load_cases(folder):
cases = []
for path in sorted(Path(folder).glob("*.yaml")):
cases.extend(yaml.safe_load(path.read_text()))
return cases
def expand(cases):
"""One param per (case, variant). Variant 0 is the clean spoken text."""
for case in cases:
n = max(len(t.get("stt_variants", [])) for t in case["turns"]) + 1
for v in range(n):
texts = [([t["say"]] + t.get("stt_variants", []))[min(v, len(t.get("stt_variants", [])))]
for t in case["turns"]]
yield case, v, texts
@dataclass
class Call:
turn: int
name: str
args: dict
@dataclass
class Verdict:
case_id: str
failures: list = field(default_factory=list)
exact_args: bool = True
@property
def passed(self): return not self.failures
def _arg_ok(actual, spec):
fn = NORMALIZERS[spec.get("match", "exact")]
try:
return fn(actual) == fn(spec["value"])
except Exception:
return False
def score(case, calls, said):
exp, v = case["expect"], Verdict(case["id"])
names = [c.name for c in calls]
for bad in exp.get("forbidden_tools", []):
if bad in names:
v.failures.append(f"forbidden tool called: {bad}")
for want in exp.get("tools", []):
hits = [c for c in calls if c.name == want["name"]]
if not hits:
v.failures.append(f"missing call: {want['name']}")
continue
last = hits[-1] # judge the final attempt; retries are capped by max_calls
for arg, spec in want.get("args", {}).items():
if arg not in last.args:
v.failures.append(f"{want['name']}.{arg} missing")
elif not _arg_ok(last.args[arg], spec):
v.failures.append(f"{want['name']}.{arg}={last.args[arg]!r} != {spec['value']!r}")
elif last.args[arg] != spec["value"]:
v.exact_args = False # passed only after normalization
if "not_before_turn" in want and hits[0].turn < want["not_before_turn"]:
v.failures.append(f"{want['name']} fired before the confirmation turn")
for name, cap in exp.get("max_calls", {}).items():
if names.count(name) > cap:
v.failures.append(f"{name} called {names.count(name)}x (cap {cap})")
if exp.get("no_tool_calls") and calls:
v.failures.append(f"expected a clarifying question, got {names}")
text = norm_text(" ".join(t for _, t in said))
spoken = exp.get("spoken", {})
if spoken.get("must_include_any") and not any(norm_text(s) in text for s in spoken["must_include_any"]):
v.failures.append(f"reply lacks any of {spoken['must_include_any']}")
for s in spoken.get("must_not_include", []):
if norm_text(s) in text:
v.failures.append(f"reply contains {s!r}")
return v
def pass_hat_k(runs_by_case, k):
"""tau-bench pass^k: mean over cases of C(c, k) / C(n, k)."""
vals = [comb(sum(r), k) / comb(len(r), k) for r in runs_by_case.values() if len(r) >= k]
return sum(vals) / len(vals)The mock backend holds fixture state and injects faults. In LiveKit's `mock_tools`, a mock receives only the parameters it declares, and returning an exception makes the tool raise. Bound methods with explicit parameters satisfy both rules.
# backend.py -- fixture-driven mock (illustrative; one method per tool)
import copy, json
from tct_core import norm_alnum
class MockBackend:
def __init__(self, fixture, faults=None):
self.state = copy.deepcopy(fixture)
self.faults = faults or {}
self.writes = []
def _fault(self, tool):
kind = self.faults.get(tool)
if kind == "timeout":
return TimeoutError(f"{tool} timed out")
if kind == "error":
return RuntimeError(f"{tool} failed")
return None
def lookup_order(self, order_id: str):
if err := self._fault("lookup_order"):
return err
row = self.state.get("orders", {}).get(norm_alnum(order_id))
return json.dumps(row) if row else "NOT_FOUND"
def cancel_order(self, order_id: str):
self.writes.append(("cancel_order", order_id))
return "CANCELLED"
def tools(self):
return {"lookup_order": self.lookup_order, "cancel_order": self.cancel_order}The test runs each case and variant, reads tool calls from the raw events (not from what the mock received), and writes one JSON line per run for pass^k.
# test_tool_cases.py -- LiveKit Agents adapter (illustrative)
import json, os
import pytest
from livekit.agents import AgentSession, mock_tools
from agent import SupportAgent # your agent; accepts a pinned clock
from backend import MockBackend
from tct_core import Call, expand, load_cases, score
CASES = list(expand(load_cases("cases")))
REPEATS = int(os.getenv("TCT_REPEATS", "1"))
@pytest.mark.asyncio
@pytest.mark.parametrize("rep", range(REPEATS))
@pytest.mark.parametrize("case,variant,texts", CASES,
ids=[f"{c['id']}-v{v}" for c, v, _ in CASES])
async def test_tool_case(case, variant, texts, rep):
fx = case["fixture"]
backend = MockBackend(fx, faults=case.get("faults"))
calls, said = [], []
async with AgentSession() as session: # uses the agent's own LLM
with mock_tools(SupportAgent, backend.tools()):
await session.start(SupportAgent(now=fx["clock"], caller=fx.get("caller")))
for turn, text in enumerate(texts):
result = await session.run(user_input=text)
for ev in result.events:
if ev.type == "function_call":
calls.append(Call(turn, ev.item.name, json.loads(ev.item.arguments or "{}")))
elif ev.type == "message" and ev.item.role == "assistant":
said.append((turn, ev.item.text_content or ""))
verdict = score(case, calls, said)
with open("tct_results.jsonl", "a") as f:
f.write(json.dumps({"id": case["id"], "variant": variant, "rep": rep,
"passed": verdict.passed, "exact_args": verdict.exact_args,
"failures": verdict.failures, "tags": case.get("tags", [])}) + "\n")
assert verdict.passed, "\n".join(verdict.failures)Run it with `TCT_REPEATS=4 pytest -q`, then group `tct_results.jsonl` by case ID and call `pass_hat_k(runs, 4)`. For Pipecat, keep `tct_core.py` and `backend.py` and swap the adapter for one that drives your pipeline in text and collects `FunctionCallParams.arguments` from your handlers. The LiveKit testing guide and Pipecat testing guide cover the surrounding setup.
Three details matter more than they look. Pin the clock: an agent that reads the system time cannot pass APT-02 deterministically. Read arguments from events: a mock that forgets to declare a parameter silently drops it, so its view of the call is incomplete. Separate text and audio tiers: text mode tests the LLM against STT renderings cheaply, while real speech through your STT is the only way to know which renderings actually occur.
How to test tool calling in your voice agent
1. List every tool and tag the writes. Mark each tool as read, reversible write or irreversible write. Irreversible writes get `not_before_turn` confirmation and `max_calls: 1` in every case that touches them.
2. Copy the use-case tables that match your agent. Rename tools to yours. Delete cases that do not apply, and keep the forbidden-tool cases even when they seem obvious.
3. Write fixtures with decoys and fresh values. Put a near-miss record one letter or digit away from each target. Never reuse values from your prompt or tool descriptions.
4. Collect real STT renderings. Run 20 to 30 real or recorded utterances of your entity types through your production STT config and paste the outputs into `stt_variants`. Include the bad ones.
5. Pin the clock and the caller. Every fixture gets a clock with a UTC offset and a caller record, including time zone and verification state.
6. Run text mode on every pull request. Use 3 to 5 repeats and fail the build on any forbidden-tool violation, confirmation miss or duplicate write.
7. Run audio mode nightly. Play the same cases as speech through your real STT at 8 kHz, with noise and at least two accents. Compare against text mode to see which failures STT introduces.
8. Gate releases on pass^k, sliced by tag. Report pass^4 for `irreversible`, `money`, `date` and `alphanumeric` separately. A single blended score hides the slice that hurts callers.
9. Feed production failures back. Every real call where a tool fired wrong becomes a new case with its real transcript as a variant. Store them in your golden dataset.
Where independent evaluation fits
Your team can run this library in CI. That covers regressions you know about. What it does not cover is the gap between your test inputs and real callers: the STT renderings you have not seen, the accents not in your team, and the model update that changes argument formatting overnight, which the LLM update regression guide covers.
Evalgent is an independent evaluator. Teams use it in three places: a pre-launch audit that runs cases like these with human-recorded speech over real telephony, a bake-off when choosing between STT or LLM vendors on the same cases, and regression scoring after each release or provider change. Results are scored per case, so they map back to the IDs in your YAML. If a staging environment is part of your release path, that is where these runs belong.
Frequently asked questions
What are voice agent tool calling test cases?
They are scripted calls with a pinned backend state. Each one feeds caller utterances, including realistic STT renderings, and asserts which tool fires, with which arguments after normalization, which tools must never fire, and what the agent says back. They differ from text tests because they include the errors speech recognition introduces before the LLM sees anything.
How do I test function calling in a voice agent without real phone calls?
Run the agent in text mode with a mocked backend, feeding both the clean utterance and the variants your STT actually produces. LiveKit's unit-test API and mock tools support this directly. Then run a smaller nightly set as real audio through your STT at 8 kHz to confirm which variants occur in practice.
How many test cases does a voice agent need?
Start with 8 to 12 per tool-heavy use case, focused on the failure modes in this library: entities, relative dates, irreversible actions, timeouts, retries, partial information and barge-in. Multiply by STT variants and audio conditions rather than writing hundreds of near-duplicate cases. Add every production failure as a new case.
Should I compare tool arguments with exact match or normalized match?
Gate on normalized match and track exact match separately. Normalized match (case, separators, currency format, ISO date) tells you whether the meaning is right. The gap between the two tells you whether your tool layer must normalize inputs itself, which it should for any ID, amount or date.
What is pass^k and why use it for voice agents?
pass^k, introduced by tau-bench, is the probability that a case passes in all k independent runs. Callers repeat the same tasks, so consistency matters more than a single success. A case that passes 95% of the time has pass^5 of about 0.77, which a single-run accuracy number hides.
How do I test that a collections agent never takes payment after a dispute?
Write a multi-turn case where the caller agrees to pay, then disputes the debt. List every payment and promise-to-pay tool under `forbidden_tools` from the dispute turn on, and assert that the dispute tool fires. Have counsel define the exact rule; the test only encodes it.
How should a voice agent handle a tool timeout?
It should say plainly that the action did not complete or could not be confirmed, never invent a result, and check status before retrying any write. Test with a fault that times out after the backend commits, then assert exactly one write in backend state and a truthful spoken reply.
How do I test barge-in during a tool call?
Script a second user turn that arrives while a slow mocked tool is running, such as a correction or an added detail. Assert one write, not two, and that the agent reports what actually happened. Check your framework's defaults, since interruptions can cancel a tool handler after the backend has already acted.
The bottom line
Tool calling in a voice agent fails at the argument, not the tool, and the argument fails in predictable places: letters that share a vowel, numbers that share a stem, dates relative to a "now" nobody pinned, and writes repeated after a timeout. Copy the cases that match your agent, add your real STT renderings and decoy records, gate on forbidden tools and pass^k, and every release will tell you which callers it would have failed.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more