Evalgent
Back to Blog
Voice AI Testing

Voice Agent Test Cases vs Assertions: What Each Is, Examples, and How to Version Them

Deepesh Jayal
19 min read
Voice Agent Test Cases vs Assertions: What Each Is, Examples, and How to Version Them
On this page

Most teams that build a voice agent in-house start testing the same way. Someone writes a list of "test calls" in a spreadsheet, dials the agent, and marks each row pass or fail. Each row mixes three things: what the caller does, what should happen, and a verdict.

That works for ten rows. It stops working when a prompt change fixes one row and quietly breaks something else in a row that still says "pass." It also stops working when nobody can say which prompt, which STT model, or which LLM snapshot produced last Tuesday's results.

The fix is to pull apart two ideas the spreadsheet blends together: the test case (the situation) and the assertion (the check). This post defines both, works through a voice example, gives you a table of assertion types with their failure modes, shows how the same assertion code becomes a production monitor, and lays out a repo and run manifest so prompts, cases, assertions and results version together. The Python at the end runs.

84% to 51%
GPT-4 accuracy on the same prime-number task, March vs June 2023 (Chen, Zaharia, Zou)
80 of 1,000
Unique completions at temperature 0 from one prompt on a standard inference stack (Thinking Machines)
under 25%
pass^8 for a strong function-calling agent on tau-bench retail (Yao et al.)

Test case vs assertion: the definitions

A test case describes a situation you put the agent in. It answers "what happens on this call?" It has four parts:

  • Setup and fixtures. Backend state before the call: the caller's account, existing appointments, the clock, which tools return what. Fixtures make the expected outcome knowable.
  • Caller persona. Who is calling and how they behave: impatient, elderly, non-native speaker, someone who corrects themselves mid-sentence.
  • Utterances or audio. Either scripted lines a simulated caller says (sent through TTS), or recorded WAV files.
  • Conditions. The channel the call travels through: 8 kHz telephony codec, background noise at a set SNR, packet loss, a barge-in at a set time.

An assertion is a single checkable expectation about what happened. It answers "did this one thing go right?" It returns pass, fail, or not applicable, and it has a severity. "The agent booked Thursday" is an assertion. "The agent never said the caller's full card number" is another.

The relationship is one to many. A case is the stimulus. Assertions are the probes you attach to the response. A failed case tells you something went wrong. A failed assertion tells you what.

Test caseAssertion
AnswersWhat situation is the agent in?Did one specific thing go right?
ContainsFixtures, persona, utterances/audio, conditionsA check function, parameters, severity
OutputA call trace (turns, timings, tool calls, backend state)Pass, fail or not applicable, with a reason
Changes whenThe product adds a flow or you find a new failureYour definition of "correct" changes
Runs on live calls?No, live calls have no fixtureMany can, as monitors
Typical countDozens to hundreds3 to 10 per case

That last-but-one row is the reason the split matters. Test cases only exist in your harness. Assertions, if written against a common trace format, can run on every production call too. More on that below.

One case, many assertions: a worked voice example

Here is a rescheduling case where the caller corrects the day halfway through. It uses the same fixture-and-turns shape as the tool-calling test case library, so read that post for the full schema. This one focuses on the assertion list.

# cases/scheduling/sched-014.yaml
id: sched-014
title: Reschedule, caller corrects the day mid-sentence
tags: [scheduling, self-correction, telephony_8k]
fixture:
  clock: "2026-10-05T09:12:00-05:00"
  caller: {phone: "+13125550142", verified: true}
  appointments:
    A-7731: {date: "2026-10-06", time: "14:00", provider: "Dr. Ruiz"}
persona: "Busy parent, speaks fast, changes mind once, does not repeat unless asked"
turns:
  - say: "Hi, I need to move my appointment with Dr. Ruiz."
  - say: "Can we do Tuesday, no wait, Thursday afternoon?"
  - say: "Two is fine."
conditions: {codec: g711_ulaw, noise: {type: kitchen, snr_db: 15}}
assertions:
  - {id: disclosure_5s, type: said_within, pattern: "recorded|virtual assistant", within_ms: 5000, severity: critical}
  - {id: reschedule_args, type: tool_called, name: reschedule_appointment,
     args: {appointment_id: "A-7731", new_date: "2026-10-08"}, severity: critical}
  - {id: no_cancel, type: tool_not_called, name: cancel_appointment, severity: critical}
  - {id: readback, type: readback_matches_tool, tool: reschedule_appointment, arg: new_date, severity: major}
  - {id: backend_state, type: state_equals, path: appointments.A-7731.date, value: "2026-10-08", severity: critical}
  - {id: overlap, type: max_overlap_ms, limit: 600, severity: minor}
  - {id: tone, type: judged, rubric: rubrics/polite-v3.md, severity: minor}

One run of this case produces one trace. Seven assertions read that trace. Suppose the agent books Tuesday instead of Thursday. `reschedule_args` and `backend_state` fail. `readback` fails too if the agent read back Thursday but sent Tuesday to the tool. The other assertions pass. That pattern is a diagnosis: the model latched onto the first day it heard, and the confirmation step did not catch it.

If the agent books Thursday but speaks over the caller for 900 ms while doing it, only `overlap` fails. Same case, different defect, and a minor one. A single pass/fail column would have hidden both distinctions.

One voice agent test case with fixtures, persona, utterances and line conditions produces a single call trace that seven assertions check, each returning pass or fail with severity

Assertion types: where each one breaks

Not all assertions are equally trustworthy. Some are brittle by design. Some need a fixture and can never leave your harness. This table covers the eight types a voice agent needs.

TypeVoice exampleBrittle whenRobust whenHow to score
Exact or normalized textAgent says the confirmation code "BD4471"Matching raw TTS input word for word; STT formatting of numbers changesNormalize case, punctuation and number words before comparing; assert on a short required span, not a whole sentenceBinary
Entity matchDate, amount, order ID captured correctlyComparing "Thursday" to "2026-10-08" as stringsResolve to a canonical value (ISO date, cents, uppercase alphanumeric) firstBinary per entity
Tool call and arguments`reschedule_appointment(A-7731, 2026-10-08)`Asserting exact call order when order does not matter; asserting optional argsAssert the required args with normalization, plus a list of tools that must not fireBinary; also count extra calls
State or backend effectAppointment row now says October 8Fixture is shared between parallel runsFresh fixture per run; read state after the call endsBinary
Timing and latencyRecording disclosure starts within 5 s of answerMeasuring from the wrong event (SIP 200 OK vs first RTP packet vs first agent audio)Define the start and end events explicitly in the assertionThreshold; also log the raw ms
Audio-levelNo agent/caller overlap longer than 600 msUsing LLM output text as a proxy for what was playedCompute from the two channel timelines of the recordingThreshold on worst case per call
Semantic or judged"Agent was polite and did not blame the caller"Vague rubric, no escape option, judge model swapped silentlyNarrow rubric, a "cannot tell" option, judge model and rubric versionedPass, fail or abstain; track agreement with human labels
Safety and policyNever reads back a full card number; never takes payment after a disputeRegex only checks one formatting of the numberCheck the tool log and the spoken text; test spaced, grouped and spelled-out digitsBinary, always critical

Three notes on the table.

Prefer the most mechanical check that answers the question. If you can assert on the tool log, do not ask a judge whether "the appointment was booked." Tool logs do not hallucinate. Judges can. When you do need a judge, keep the rubric narrow and give it a way out. Zheng et al. found strong LLM judges reached over 80% agreement with human preferences, about the same rate humans agree with each other, and they also documented position, verbosity and self-enhancement biases. Treat an abstain as "not applicable," never as a pass. If you judge many calls on narrow classifications, a dedicated classifier can be cheaper and more repeatable than a general LLM; see how Jev's accuracy is measured.

Your criteria will move after you read outputs. Shankar et al. call this criteria drift: people need criteria to grade outputs, but grading outputs changes the criteria. Practically, the first 50 transcripts you read will rewrite your rubric. That is fine, as long as the rubric is versioned and old results are marked as graded under the old rubric.

Your prompt history is an assertion backlog. The SPADE paper observed that developers add instructions to prompts each time they find a bad output, and it mined those prompt edits into candidate assertions. Across nine real pipelines it cut the number of assertions by 14% and false failures by 21% versus simpler baselines (Shankar et al.). The voice version is simple: every time you add a line like "always read the date back before booking" to the system prompt, add the assertion that checks it in the same commit.

Severity, thresholds and repeat runs

LLM-driven agents do not give the same answer twice, even with temperature set to zero. Atil et al. ran five models configured to be deterministic across eight tasks and saw accuracy vary by up to 15% between runs. Thinking Machines sampled one prompt 1,000 times at temperature 0 and got 80 unique completions, and traced the cause to inference kernels that are not batch-invariant. You do not control the batch your request lands in on a hosted API. So a voice test case has to run more than once, and a seed does not make a hosted LLM repeatable.

pass^k, with numbers

Tau-bench introduced pass^k: the probability that an agent succeeds on all k independent attempts at the same task. With n runs and c successes, the unbiased estimate is C(c, k) / C(n, k). The paper reported a strong function-calling agent at under 25% pass^8 in its retail domain, far below its single-try success rate.

Worked example. A case passes 9 of 10 runs.

  • pass^1 = 9/10 = 0.90
  • pass^3 = C(9,3)/C(10,3) = 84/120 = 0.70
  • pass^5 = C(9,5)/C(10,5) = 126/252 = 0.50

A 90% case looks fine on a dashboard. Measured as "works every time a caller hits it across five calls," it is a coin flip.

Why statistics cannot gate your critical assertions

Here is a result most teams find surprising. Say assertion `reschedule_args` passed 10 of 10 runs on the old prompt. How far must it drop on the new prompt before a one-sided Fisher exact test says it is a real regression at p < 0.05?

New prompt passesp-valueVerdict
9/100.50Noise
8/100.24Noise
7/100.105Noise
6/100.043Regression

Even at 20 runs per side, 20/20 dropping to 17/20 gives p = 0.115. The flip side is the rule of three: zero failures in n runs only bounds the true failure rate below about 3/n at 95% confidence. Ten clean runs still allow a 30% failure rate. Sixty clean runs are needed to claim under 5%.

So use two different gates:

  • Critical (safety, compliance, money movement, wrong booking): zero tolerance. Any failure in any repeat blocks release. Statistics are not involved; one failure is an existence proof.
  • Major (task outcome, readback, correct routing): block on a statistically significant drop, and track pass^k per case over time.
  • Minor (tone, small overlaps, soft latency): trend only. Never blocks a release on its own.

These tiers are a starting point you should tune to your product, not a standard. The practical rule underneath them is fixed, though: put the most repeats on the cases whose assertions are critical, because those are the ones where one failure matters. The prompt-change release workflow covers how to wire these gates into deploys.

Assertions that also run on production calls

A test case needs a fixture to know the right answer. A live call has no fixture. That is why "how do assertions generalize from tests to production?" has a precise answer: some do, some cannot, and you can tell which by asking what the assertion needs to know.

Sort every assertion into one of four classes:

1. Invariants. True on every call regardless of context. Disclosure within 5 s, never speaking a full card number, overlap under 600 ms, no payment tool after the word "dispute." These run unchanged on production traces.

2. Self-referential checks. They compare the call to itself. "The date read back to the caller equals the date sent to the booking tool." "What the agent told the caller matches the backend row after the call." No fixture needed, because the call supplies both sides. These are the most valuable production monitors.

3. Fixture-bound checks. They need the known right answer: "booked 2026-10-08 for A-7731." These stay in the harness.

4. Judged checks. They run on both, but in production you sample, and you track the abstain rate.

The bridge is a shared trace format. If your harness and your production logging both emit the same structure (turns with role, text and start/end ms; tool calls with args and timestamps; backend state after the call; overlap intervals from the recording), the same `evaluate()` function runs on both. A failed monitor on a live call then points to an assertion ID you already understand.

Map sorting voice agent assertions into invariant, self-referential, fixture-bound and judged classes, showing which also run as production monitors and how failed live calls become test cases

Two gotchas when you feed production traces to test-built assertions:

  • The agent's text is not what the caller heard. When a caller barges in, the LLM may have generated a full sentence while the TTS played only half. Assertions about what was "said" must use the played-out portion, cut at the interruption timestamp, or a `never_said` check will fire on words that were never spoken.
  • The caller's text is an STT guess. A study of Whisper found roughly 1% of transcriptions contained whole hallucinated phrases that were not in the audio, and hallucinations were more common for speakers with long pauses. Assertions about the caller's words inherit those errors. Assertions about tool logs and backend state do not. Weight your monitors toward the latter.

For the monitoring side of this, including sampling and alerting, see monitoring voice agents in production.

Turning failed production calls into regression test cases

Every failed monitor is a candidate test case. Converting it well takes six steps, and two of them are where most teams lose the bug.

1. Capture the version that served the call. At session start, write the run manifest ID (or at least the prompt hash and model IDs) into the call's metadata. Without it you cannot reproduce the failure on the same configuration.

2. Pull the full trace and both audio channels. Transcripts alone lose timing, overlaps and the exact audio the STT heard.

3. Redact with format-preserving surrogates. Replace a 10-digit account number with a different 10-digit number, a card number with a Luhn-valid test number, a name with a name of similar length. Replacing everything with `[REDACTED]` often deletes the bug, because many voice bugs are about the format itself: a 10-digit number heard as 9, a name that sounds like a common word. For audio, either re-synthesize the caller's turns from redacted text with TTS, or mute the sensitive spans using word timestamps. Re-synthesis loses accent and noise; muting keeps them.

4. Minimize. A 14-turn call that fails at turn 11 rarely needs all 14 turns. Delta debugging, from Zeller and Hildebrandt, removes chunks of input while the failure still reproduces. In their case study it reduced a browser crash from 95 user actions to 3. For voice, drop turns or halves of turns, rerun, keep the shortest version that still fails. Shorter cases run faster and fail for one reason.

5. Write the oracle. Add the fixture that makes the right answer knowable and the assertion that failed, at critical or major severity.

6. Confirm red, then green. Run the new case on the production version and confirm it fails. Run it on the fix and confirm it passes. A case that never failed proves nothing.

Tag the case with provenance: a hash of the source call ID, the date, and the manifest that served the call. The production feedback loop covers triage and prioritization, and the golden dataset guide covers when a promoted case should join the frozen benchmark set.

Versioning prompts, test cases, assertions and results together

Results mean nothing without the versions that produced them. There are at least nine moving parts behind one voice agent test result: the system prompt, the flow or tool schema, the test case, its fixtures and audio, the assertion code, the judge rubric, the STT model, the LLM, and the TTS voice. Change any one and the result is a different experiment.

Hosted models move under you. Chen, Zaharia and Zou measured GPT-4 identifying prime numbers at 84% accuracy in March 2023 and 51% in June 2023 on the same questions. STT providers do the same. Deepgram's `/v1/listen` takes a `version` parameter that defaults to `latest`, and every response includes `metadata.model_info` with the model `name`, `arch` and `version` string, plus a `request_id`. Record that `version` string in your manifest on every run, because "nova-3" alone does not identify the model that transcribed your audio. The LLM update regression guide covers pinning on the LLM side.

A repo layout that keeps them together

voice-agent/
  prompts/              system.md, tools.md (what ships)
  flows/                state machine or tool schemas (JSON Schema)
  cases/
    scheduling/sched-014.yaml
    billing/bill-003.yaml
    from_prod/          promoted production failures, with provenance
  assertions/
    vassert.py          assertion library (code, versioned in git)
    rubrics/polite-v3.md  judge rubrics (versioned like code)
  fixtures/
    audio/              WAV/FLAC + .sha256 sidecar per file
    backends/           seed data per case
  config/
    models.yaml         stt/llm/tts IDs and versions, temperature, endpointing
  results/              gitignored; synced to object storage
    <run_id>/manifest.json
    <run_id>/results.jsonl
  runs.csv              committed: run_id, date, git_sha, summary pass rates

Prompts, flows, cases, assertions and rubrics live in git, in one repo, so one commit pins all of them. Audio fixtures are large, so store them in object storage or Git LFS, keyed by SHA-256, and commit only the hash. Results go to object storage under a run ID, with a one-line summary committed to `runs.csv` so the history is reviewable in a pull request.

The run manifest

The manifest is a small JSON file written with every run. It answers "exactly what produced these results?"

  • `git_sha` and `git_dirty` (a dirty tree means the SHA does not describe the code)
  • `prompt_hash`, `flow_hash`, `cases_hash`, `fixtures_hash`: content hashes, so an uncommitted edit still changes the ID
  • `assertion_code`: a hash of each assertion function's source, plus rubric hashes and the judge model ID
  • `models`: STT, LLM and TTS model IDs with the provider-reported versions
  • `config_hash`: temperature, endpointing and VAD settings, voice ID
  • `repeats` and `seed`

Hash the manifest itself and use that as the run ID. That makes results content-addressed: the same inputs always map to the same directory, so a rerun with nothing changed is detected and skipped, and two runs with different IDs are guaranteed to differ in some recorded input.

Comparing runs: diff per assertion, not per suite

A suite-level pass rate going from 91% to 90% says nothing. Compare at the level of (case, assertion) and classify each pair:

  • Regressed: pass rate fell, and either the assertion is critical or the drop is significant.
  • Worse, within noise: pass rate fell, but not significantly. Rerun with more repeats if it is major.
  • Improved, added, removed.
  • Not comparable: the assertion's code hash changed between runs.

That last category prevents a quiet mistake. If you loosen a regex in an assertion, the pass rate rises, and without the code hash it looks like the agent got better. For the side-by-side view across prompt versions, the prompt comparison guide covers how to present it.

The same trick answers "how do I compare production logs across prompt versions?" Because every live call carries its manifest ID in metadata, you can group production monitor results by `prompt_hash` and run the same per-assertion diff on live traffic, with the caveat that live traffic mix differs between periods.

Portable formats so cases are not locked to a vendor

Keep cases in YAML (human-edited) or JSONL (generated), audio as WAV or FLAC with hashes, and assertions as code you own. Any test platform, including the one you might switch to, should consume those through an adapter. If you are leaving a platform, ask for the export in this order: case definitions including persona prompts, assertion definitions including judge rubrics (names alone are useless), the audio, and historical results with the versions that produced them.

Two voice agent run manifests compared side by side, with a per-assertion diff grid marking each pair as regressed, within noise, improved or not comparable

The code: assertion library, runner with manifest, run diff

This is illustrative and simplified, but it runs on Python 3.10+ with PyYAML. You supply `run_agent(case, seed)`, which drives your agent (a simulated caller through LiveKit or Pipecat, or text-only for fast checks) and returns a trace in the shared format. For fast, no-phone-call harnesses, see replacing manual test calls.

vassert.py: a tiny assertion library

# vassert.py (simplified)
import hashlib, inspect, re

REGISTRY = {}

def assertion(kind):
    """Register an assertion type and fingerprint its source code."""
    def wrap(fn):
        fn.code_hash = hashlib.sha256(inspect.getsource(fn).encode()).hexdigest()[:12]
        REGISTRY[kind] = fn
        return fn
    return wrap

def _norm(s):
    return re.sub(r"[^a-z0-9]", "", str(s).lower())

@assertion("tool_called")
def tool_called(trace, name, args=None):
    calls = [c for c in trace["tool_calls"] if c["name"] == name]
    if not calls:
        return False, f"{name} never called"
    want = {k: _norm(v) for k, v in (args or {}).items()}
    for c in calls:
        if {k: _norm(c["args"].get(k)) for k in want} == want:
            return True, ""
    return False, f"args {calls[-1]['args']} != {args}"

@assertion("tool_not_called")
def tool_not_called(trace, name):
    hit = [c for c in trace["tool_calls"] if c["name"] == name]
    return (not hit), (f"{name} called {len(hit)}x" if hit else "")

@assertion("said_within")
def said_within(trace, pattern, within_ms):
    for t in trace["turns"]:
        if t["role"] == "agent" and re.search(pattern, t["text"], re.I):
            return t["start_ms"] <= within_ms, f"first match at {t['start_ms']} ms"
    return False, "never said"

@assertion("never_said")
def never_said(trace, pattern):
    for t in trace["turns"]:  # use played-out text, cut at barge-in
        if t["role"] == "agent" and re.search(pattern, t["text"], re.I):
            return False, f"said at {t['start_ms']} ms"
    return True, ""

@assertion("max_overlap_ms")
def max_overlap_ms(trace, limit):
    worst = max(trace.get("overlaps_ms") or [0])
    return worst <= limit, f"worst overlap {worst} ms"

@assertion("state_equals")
def state_equals(trace, path, value):
    cur = trace["backend_after"]
    for key in path.split("."):
        cur = cur.get(key, {}) if isinstance(cur, dict) else None
    return _norm(cur) == _norm(value), f"{path}={cur!r}"

@assertion("readback_matches_tool")
def readback_matches_tool(trace, tool, arg):
    """Self-referential: needs no fixture, so it also runs on live calls."""
    calls = [c for c in trace["tool_calls"] if c["name"] == tool]
    if not calls:
        return None, "not applicable"
    val = _norm(calls[-1]["args"].get(arg, ""))
    spoken = _norm(" ".join(t["text"] for t in trace["turns"]
                   if t["role"] == "agent" and t["start_ms"] < calls[-1]["t_ms"]))
    return (val in spoken), f"{arg}={val!r} read back: {val in spoken}"

def evaluate(trace, spec):
    fn = REGISTRY[spec["type"]]
    params = {k: v for k, v in spec.items() if k not in ("id", "type", "severity")}
    passed, detail = fn(trace, **params)
    return {"assertion_id": spec["id"], "type": spec["type"],
            "code_hash": fn.code_hash, "severity": spec.get("severity", "major"),
            "passed": passed, "detail": detail}

In a real suite, `readback_matches_tool` needs a date normalizer so "Thursday, October 8th" matches "2026-10-08". The judged type is left out here; it would call your judge, return `None` on abstain, and include the rubric hash and judge model ID in its `code_hash`.

run_suite.py: runner that writes a run manifest

# run_suite.py (simplified)
import glob, hashlib, json, pathlib, subprocess, time, yaml
from vassert import evaluate, REGISTRY

def sha(data: bytes) -> str:
    return hashlib.sha256(data).hexdigest()[:16]

def file_hash(pattern):
    h = hashlib.sha256()
    for p in sorted(glob.glob(pattern, recursive=True)):
        h.update(p.encode()); h.update(pathlib.Path(p).read_bytes())
    return h.hexdigest()[:16]

def git(*args):
    try:
        return subprocess.run(["git", *args], capture_output=True, text=True).stdout.strip()
    except FileNotFoundError:
        return ""

def build_manifest(config, repeats, seed):
    return {
        "git_sha": git("rev-parse", "HEAD") or "unknown",
        "git_dirty": bool(git("status", "--porcelain")),
        "prompt_hash": file_hash("prompts/**/*"),
        "flow_hash": file_hash("flows/**/*"),
        "cases_hash": file_hash("cases/**/*.yaml"),
        "fixtures_hash": file_hash("fixtures/**/*"),
        "assertion_code": {k: f.code_hash for k, f in sorted(REGISTRY.items())},
        "models": config["models"],  # include provider-reported versions
        "config_hash": sha(json.dumps(config, sort_keys=True).encode()),
        "repeats": repeats, "seed": seed,
    }

def run(run_agent, config, repeats=5, seed=7):
    manifest = build_manifest(config, repeats, seed)
    run_id = sha(json.dumps(manifest, sort_keys=True).encode())
    out = pathlib.Path("results") / run_id
    if (out / "results.jsonl").exists():
        print(f"run {run_id} exists: same inputs, skipping")
        return run_id
    out.mkdir(parents=True, exist_ok=True)
    manifest["started_at"] = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
    with open(out / "results.jsonl", "w") as f:
        for path in sorted(glob.glob("cases/**/*.yaml", recursive=True)):
            case = yaml.safe_load(open(path))
            for i in range(repeats):
                trace = run_agent(case, seed=seed + i)
                for spec in case["assertions"]:
                    row = evaluate(trace, spec)
                    row.update(case_id=case["id"], repeat=i)
                    f.write(json.dumps(row) + "\n")
    (out / "manifest.json").write_text(json.dumps(manifest, indent=2))
    return run_id

The `started_at` timestamp is added after hashing on purpose. If it were inside the hash, every run would get a new ID and content addressing would stop working.

diff_runs.py: per-assertion diff of two runs

# diff_runs.py (simplified)
import json, sys
from collections import defaultdict
from math import comb

def load(run_id):
    agg = defaultdict(lambda: {"n": 0, "pass": 0, "hash": None, "sev": None})
    for line in open(f"results/{run_id}/results.jsonl"):
        r = json.loads(line)
        if r["passed"] is None:  # not applicable or judge abstained
            continue
        a = agg[(r["case_id"], r["assertion_id"])]
        a["n"] += 1; a["pass"] += bool(r["passed"])
        a["hash"], a["sev"] = r["code_hash"], r["severity"]
    return agg

def p_worse(p_a, n_a, p_b, n_b):
    """One-sided Fisher exact: chance B scores this low if A and B are equal."""
    total, passes = n_a + n_b, p_a + p_b
    lo = max(0, passes - n_a)
    return sum(comb(n_b, k) * comb(n_a, passes - k)
               for k in range(lo, p_b + 1)) / comb(total, passes)

def diff(base, cand, alpha=0.05):
    A, B = load(base), load(cand)
    for key in sorted(set(A) | set(B)):
        a, b = A.get(key), B.get(key)
        if a is None or b is None:
            status = "added" if a is None else "removed"
        elif a["hash"] != b["hash"]:
            status = "NOT COMPARABLE (assertion code changed)"
        elif b["pass"] / b["n"] < a["pass"] / a["n"]:
            p = p_worse(a["pass"], a["n"], b["pass"], b["n"])
            critical = b["sev"] == "critical"
            status = "REGRESSED" if (critical or p < alpha) else f"worse, within noise (p={p:.2f})"
        elif b["pass"] / b["n"] > a["pass"] / a["n"]:
            status = "improved"
        else:
            continue
        fmt = lambda x: f"{x['pass']}/{x['n']}" if x else "-"
        print(f"{key[0]:<12} {key[1]:<18} {(b or a)['sev']:<9} "
              f"{fmt(a):>6} -> {fmt(b):<6} {status}")

if __name__ == "__main__":
    diff(sys.argv[1], sys.argv[2])

Run against a simulated agent that books the wrong day 6 times out of 10 on a new prompt, the output looks like this:

sched-014    backend_state      critical   10/10 -> 4/10   REGRESSED
sched-014    reschedule_args    critical   10/10 -> 4/10   REGRESSED

The readback assertion still passes in that run, because the agent read back the wrong date and then booked it. That is the self-referential check working as designed: it confirms consistency, not correctness. Correctness needs the fixture-bound assertions next to it.

How to set up versioned test cases and assertions for your voice agent

1. Split your spreadsheet. For each row, move the situation (account state, what the caller says, line conditions) into a case file, and each expected outcome into a separate assertion with a severity.

2. Define one trace format. Turns with role, text and start/end ms; tool calls with args and timestamps; backend state after the call; overlap intervals from the recording. Make both your harness and your production logging emit it.

3. Write assertions as code in your repo. Start with tool-called, tool-not-called, state-equals, said-within and never-said. Add judged assertions last, with a narrow rubric and an abstain option.

4. Classify each assertion as invariant, self-referential, fixture-bound or judged. Deploy the first two as production monitors.

5. Pin every model you can and record the provider-reported version of every model you cannot pin, per run.

6. Write a run manifest with git SHA, dirty flag, content hashes, assertion code hashes, model versions, config, repeats and seed. Use its hash as the run ID.

7. Run each case at least 5 times; run cases with critical assertions 10 times or more. Report pass^k per case.

8. Gate releases on zero critical failures and no significant major regression in the per-assertion diff. Treat "not comparable" rows as a prompt to rerun the baseline with the new assertion code.

9. Stamp production calls with the manifest ID that served them, and promote failed monitor hits into `cases/from_prod/` using format-preserving redaction and minimization.

For regulated flows, where many assertions are compliance invariants, the policy test case guide lists the checks worth making critical.

Where independent evaluation fits

Everything above can run in-house, and for many teams it should. Independent evaluation earns its place in four spots: a pre-launch audit by someone who did not write the cases or the prompt, a bake-off where two STT or LLM vendors must face the same cases and assertions, a regression check on each prompt or model change, and ongoing scoring of production calls against your invariants. Evalgent does that work on your own cases, with per-assertion results tied to the versions that produced them, and the cases stay in your portable format.

Frequently asked questions

What is the difference between voice agent test cases and assertions?

A test case is the scenario: fixtures, a caller persona, utterances or audio, and line conditions such as codec and noise. An assertion is one checkable expectation about the result, like a tool called with specific arguments or a disclosure spoken within 5 seconds. One case usually carries 3 to 10 assertions, each with its own severity.

How do I version voice agent prompts, test cases, assertions, and results together?

Keep prompts, flows, cases, assertion code and judge rubrics in one git repo so a commit pins them together. Write a run manifest per test run with the git SHA, content hashes, assertion code hashes, and STT, LLM and TTS model versions. Hash the manifest to get the run ID and store results under it.

How do I version and compare voice agent logs across different prompt versions?

Write the manifest ID or prompt hash into each production call's metadata at session start. Run your invariant and self-referential assertions on every logged call. Then group results by prompt hash and diff pass rates per assertion, the same way you diff two test runs. Remember that caller mix changes between periods.

How do assertions generalize from voice agent tests to production calls?

Only some do. Invariants (disclosure timing, never reading card numbers) and self-referential checks (readback matches the tool argument) need no fixture, so they run on live calls as monitors. Fixture-bound checks need a known right answer and stay in tests. Both sides must emit the same trace format.

How do I log voice agent prompt versions with each test run?

Hash the prompt files' contents, not just the filename, and record the hash in a run manifest next to the git SHA and a dirty-tree flag. Content hashes catch uncommitted edits. Also record model versions the provider reports in responses, such as the version in Deepgram's model_info metadata.

How can I turn failed production voice calls into regression test cases?

Pull the trace, both audio channels and the version that served the call. Redact with format-preserving surrogates so the bug survives. Minimize by removing turns while the failure still reproduces. Add a fixture and the failing assertion, confirm it fails on the old version and passes on the fix, then tag provenance.

How do I export voice agent test cases if I switch QA vendors?

Ask for case definitions including persona prompts, assertion definitions including judge rubrics, all audio fixtures, and historical results with the versions that produced them. Store cases as YAML or JSONL and audio as WAV or FLAC with SHA-256 hashes in your own repo, so any future tool reads them through an adapter.

How many assertions should one voice agent test case have?

Usually 3 to 10. Include at least one outcome check (tool call or backend state), the forbidden actions for that flow, and any invariants that apply. More than about 10 often means the case covers two situations and should be split, so a failure points to one cause.

The bottom line

Test cases describe the situation and assertions describe what correct looks like, and keeping them separate is what lets one assertion library check both your harness and every live call. Version all of it under a hashed run manifest, diff results per assertion, and you will always know which change broke what.

Related Articles