Evalgent
Back to Blog
Voice AI Testing

How to A/B Test Voice Agent Prompts: Offline Pairs, Live Splits, and the Sample-Size Math

Deepesh Jayal
Updated
18 min read
How to A/B Test Voice Agent Prompts: Offline Pairs, Live Splits, and the Sample-Size Math
On this page

Most teams that run an in-house voice agent on LiveKit or Pipecat "A/B test" a prompt by making ten test calls to each version and picking the one that felt better. That tells you almost nothing. A prompt change that lifts task success by three points needs thousands of live calls to confirm, and an LLM judge comparing two transcripts can flip its verdict just because you swapped their order.

This guide covers the full method for prompts and agent configs, not TTS voices. For voice testing, see our guide to A/B testing voices in Deepgram, which goes deep on listening tests and always-valid p-values. Here you get two halves. First, the offline paired A/B that catches most bad prompts before a caller hears them. Second, the online A/B on live traffic, with assignment code for both frameworks, worked sample-size math at mid-size volumes, and rules for deciding.

Offline A/B vs online A/B: two different questions

An offline A/B runs prompt A and prompt B against the same scripted scenarios with synthetic callers. It answers one question: does B break anything, and does it do better on the cases we know about? It's cheap, it's repeatable, and nobody real is exposed.

An online A/B splits real callers between A and B. It answers a different question: does B move the business outcome on the traffic we actually get? Offline can't answer that, because your scenarios are not your traffic. They're missing the caller who reads a 16-digit policy number in two breaths, and they're missing next month's weird week.

Use both, in order. The offline A/B is a gate. Only variants that pass it earn live traffic. Our offline vs online evaluation guide covers that split in general. The rest of this post is about running each half properly when the thing under test is a prompt.

Two-stage A/B pipeline for voice agent prompts: offline paired scenario runs with a position-swapped judge gate the variant before a caller-level live traffic split and decision rules

What counts as "the variant"

A prompt rarely changes alone. Define the variant as a bundle: system prompt, greeting, tool descriptions, model snapshot, temperature, and turn-handling settings. Hash it and log the hash on every call. If B also bumped the model alias, you're testing two things at once and can't say which one helped.

Change one thing per experiment when you can. When you can't, because a new prompt only works with a new tool schema, say so in the experiment record and treat the bundle as the unit.

Half one: the offline paired A/B

The offline test has one big advantage over live traffic: you can run both variants on exactly the same inputs. That's a paired design, and it removes most of the noise. A live A/B compares two different groups of callers. An offline A/B compares two answers to the same caller.

Run each scenario k times, not once

LLM agents are stochastic, so one run per scenario measures luck. The tau-bench paper by Yao et al. (2024) introduced pass^k: the chance an agent succeeds on all k independent tries of the same task. They found even strong function-calling agents succeeded on fewer than half the tasks, and pass^8 fell below 25% in their retail domain. An agent that "passes" once can fail the same scenario a third of the time.

For prompt A/B work, run each scenario k = 5 times per variant at your production temperature. Don't set temperature to 0 to "make it deterministic." That tests a configuration you don't ship, and hosted models aren't bit-exact at temperature 0 anyway. You want the spread callers will get.

A practical suite: 60 scenarios, 5 runs each, two variants, so 600 simulated calls. Check outcomes with deterministic assertions where you can (was the booking tool called with the right date, did the agent read the disclosure). Our post on test cases vs assertions shows how to write those so both variants are scored by the same checks.

Analyze the pairs, not the totals

Don't compare "A passed 241 of 300, B passed 255 of 300" with a two-proportion test. Those 300 runs are 60 scenarios repeated, so they're not independent, and the test will overstate your confidence. Evan Miller's Adding Error Bars to Evals (2024) makes this point for LLM evals: treat questions as samples from a larger population, cluster standard errors when items are related, and use paired differences when two models answer the same questions.

The method: for each scenario, compute B's pass rate minus A's pass rate. You get 60 differences. Their mean is the effect, and their standard error comes from those 60 numbers, not from 600 runs.

# offline_ab.py (illustrative): paired comparison of two prompt variants
# results: {scenario_id: {"A": [1,0,1,1,1], "B": [1,1,1,1,1]}}  1 = assertions passed
import math
from math import comb
from statistics import mean, stdev

def pass_hat_k(successes: int, n: int, k: int) -> float:
    """Unbiased estimate of pass^k from n runs (tau-bench style)."""
    return comb(successes, k) / comb(n, k) if n >= k else float("nan")

def paired_offline_ab(results: dict, k: int = 5) -> dict:
    diffs, pk_a, pk_b = [], [], []
    for runs in results.values():
        a, b = runs["A"], runs["B"]
        diffs.append(mean(b) - mean(a))
        pk_a.append(pass_hat_k(sum(a), len(a), k))
        pk_b.append(pass_hat_k(sum(b), len(b), k))
    d, se = mean(diffs), stdev(diffs) / math.sqrt(len(diffs))
    return {
        "mean_diff": d,
        "ci95": (d - 1.96 * se, d + 1.96 * se),
        "pass_k_A": mean(pk_a), "pass_k_B": mean(pk_b),
        "regressed": [s for s, r in results.items()
                      if mean(r["B"]) < mean(r["A"]) - 0.4],  # lost 2+ of 5 runs
    }

Read three things. The paired difference and its interval tell you if B is better on average. The pass^k numbers tell you if B is more reliable, which matters more on a phone line than the average does. The `regressed` list names scenarios where B lost two or more of five runs. Any regression on a compliance or payment scenario blocks the variant, whatever the average says.

Judging two transcripts: the position-bias trap

Some outcomes can't be asserted, like "which response handled the angry caller better." Teams reach for an LLM judge that sees both transcripts and picks a winner. The research says that judge is biased by order.

Zheng et al. (2023), the MT-Bench paper, swapped the order of two answers and checked if the judge's verdict held. With the default prompt, GPT-4 gave consistent verdicts in only 65.0% of cases and favored the first answer in 30.0%. GPT-3.5 was consistent 46.2% of the time. Claude-v1 was consistent 23.8% of the time and favored the first position 75.0% of the time. Wang et al. (2023) showed the bias can decide results outright: with ChatGPT as judge, Vicuna-13B beat ChatGPT on 66 of 80 queries just by changing the order.

Those models are old, and newer judges do better. But the mechanism hasn't gone away, and you won't know your judge's rate until you measure it.

Position bias of LLM judges from Zheng et al. 2023: consistency after swapping answer order and share of first-position wins for GPT-4, GPT-3.5 and Claude-v1, with the two-order judging protocol

The protocol that fixes it, from Zheng et al.: call the judge twice, A-first and B-first. Count a win only when the same variant wins in both orders. Any disagreement is a tie. Three more rules:

  • Blind the labels. Show "Response 1" and "Response 2," never "control" and "new prompt." Judges and humans both lean toward the label they expect to win.
  • Strip length tells. Zheng et al. also documented verbosity bias. If B's prompt makes the agent wordier, a judge may reward that, while a caller on the phone hates it. Score talk time on its own and keep it out of the judge's view of quality.
  • Report the tie rate. If more than about a third of pairs flip on swap, your rubric is too vague to support a decision. Rewrite it into narrower yes/no questions.

For more on where judges fail on voice calls, see the limits of LLM-as-judge. For a full side-by-side workflow without live traffic, see voice agent prompt comparison.

Half two: the online A/B on live calls

A variant that passes the offline gate gets real traffic. Four decisions make or break the online test: the randomization unit, how calls are pinned, which metrics count, and how many calls you need.

Randomize by caller, not by call

If you assign each call independently, a patient who calls Monday and again Thursday may hear prompt A and then prompt B. That's a worse experience. It also breaks the statistics. The second call depends on the first ("you told me Tuesday was open"), and repeat callers count twice, so your effective sample is smaller than your call count.

Assign by a stable caller key: phone number for inbound, account or contact ID for outbound. Hash it with an experiment-specific salt so each new experiment reshuffles callers instead of reusing the same split.

# assign.py (illustrative): stable, salted caller-level assignment
import hashlib

def assign(caller_key: str, experiment: str, weights: dict[str, int]) -> str:
    """weights like {"A": 50, "B": 50}; same caller + experiment -> same arm."""
    h = hashlib.sha256(f"{experiment}:{caller_key}".encode()).hexdigest()
    bucket = int(h[:15], 16) % 100
    acc = 0
    for arm, pct in weights.items():
        acc += pct
        if bucket < acc:
            return arm
    return next(iter(weights))

Don't randomize by room name or call SID. Both are unique per call, so they quietly turn a caller-level design back into a call-level one.

Pinning per call in LiveKit

In LiveKit, `agent_name` on `@server.rtc_session()` turns on explicit dispatch, and dispatch can carry job metadata (a string up to 512 KiB). Inbound SIP dispatch rules name their agents statically, with no traffic weights. So for prompt and config experiments, keep one agent and choose the variant inside the entrypoint:

# agent.py (simplified): one LiveKit agent, variant chosen once per call
import json
from livekit import agents, rtc
from livekit.agents import Agent, AgentServer, AgentSession, JobContext
from assign import assign
from variants import load_variant  # your code: prompt, tools, model per arm

EXPERIMENT = "greeting-v2"
server = AgentServer()

@server.rtc_session(agent_name="support-agent")
async def entry(ctx: JobContext):
    meta = json.loads(ctx.job.metadata or "{}")
    participant = await ctx.wait_for_participant()
    caller = meta.get("contact_id") or participant.attributes.get("sip.phoneNumber")
    arm = meta.get("arm") or (assign(caller, EXPERIMENT, {"A": 50, "B": 50})
                              if caller else "A")  # no key -> control, excluded
    v = load_variant(EXPERIMENT, arm)              # frozen for this call
    ctx.log_context_fields = {"experiment": EXPERIMENT, "arm": arm,
                              "variant_hash": v.hash, "in_test": bool(caller)}
    session = AgentSession(stt=v.stt, llm=v.llm, tts=v.tts)
    await session.start(room=ctx.room, agent=Agent(instructions=v.prompt, tools=v.tools))

if __name__ == "__main__":
    agents.cli.run_app(server)

Two gotchas. First, LiveKit's SIP participant docs note that `sip.phoneNumber` isn't available if `HidePhoneNumber` is set on the dispatch rule. Without a fallback, every call would land in one arm. The code sends key-less calls to control and flags them out of the analysis. Second, for outbound campaigns, decide the arm in your dialer and pass it as `arm` in dispatch metadata, so the assignment exists before the phone rings. Our prompt versioning guide covers separate agent names, deployments, and rollback.

Pinning per call in Pipecat

With Pipecat behind Twilio Media Streams, your webhook is the session router. It returns the TwiML that decides where the call's audio goes, so it's the natural place to assign. Pass the arm as a Stream ``. Twilio delivers it in the stream's start message, and Pipecat's `parse_telephony_websocket()` exposes it in `call_data["body"]`:

# router.py (simplified): Twilio voice webhook assigns the arm
from fastapi import FastAPI, Form, Response
from assign import assign

app = FastAPI()
EXPERIMENT = "greeting-v2"

@app.post("/voice")
async def voice(From: str = Form(...)):
    arm = assign(From, EXPERIMENT, {"A": 50, "B": 50})
    twiml = f"""<Response><Connect>
  <Stream url="wss://voice.example.com/ws">
    <Parameter name="experiment" value="{EXPERIMENT}"/>
    <Parameter name="arm" value="{arm}"/>
  </Stream></Connect></Response>"""
    return Response(content=twiml, media_type="application/xml")

# bot.py: read the pinned arm once, at session start
# transport_type, call_data = await parse_telephony_websocket(runner_args.websocket)
# arm = call_data["body"].get("arm", "A")
# variant = load_variant(EXPERIMENT, arm)

Whichever framework you use, write an exposure record when the session starts: call ID, caller key, experiment, arm, variant hash, timestamp. The analysis joins outcomes to that record. Never infer the arm later from logs. Our list of what to log on every call has the rest of the fields.

Check the split before you read the result

With a 50/50 split you expect roughly equal exposures. If you see 5,200 vs 4,800, run a chi-square test against the planned ratio. A sample ratio mismatch with p < 0.001 means something is broken: one arm crashing before it logs exposure, a retry path that reassigns, or a key fallback that skews. Don't analyze an experiment with a ratio mismatch. Find the bug.

Primary metric and guardrails for prompt tests

Pick one primary metric before launch. It decides the winner. Everything else is a guardrail that can veto it.

RoleMetricDefinitionWhy prompts move it
PrimaryTask successShare of calls where the caller's goal is verifiably done (booking written, order placed)Instruction clarity, tool descriptions, confirmation wording
GuardrailRe-ask rateShare of agent turns that repeat a question already answeredLong prompts drop context rules; wording loops
GuardrailEscalation rateShare of calls transferred to a humanPrompt tone and refusal boundaries
GuardrailAverage handle timeMean connected seconds per callWordier prompts make wordier agents
GuardrailPolicy violationsShare of calls failing a required disclosure or forbidden-action checkNew wording can bury a rule
GuardrailTime to first audio p95Caller end of speech to first agent audio, 95th percentilePrompt length, prompt-cache misses

Task success is the outcome that matters, but it's a lagging, noisy signal. Re-ask rate and violations move faster. Our post on leading and lagging voice agent metrics explains how to pair them. For the definition of the primary, see task success rate.

The latency guardrail is a cache guardrail

A longer prompt costs less latency than most teams fear. OpenAI's latency guide says halving a prompt may cut latency by only 1 to 5%. The bigger effect comes from the prompt cache, which is prefix-based. OpenAI's prompt caching docs say a cache hit requires the whole prefix to match.

That creates an A/B artifact. Your new variant has a new prefix, so it starts cold. At a 90/10 split, the 10% arm keeps its cache warm less often than the 90% arm. B looks slower on time to first audio, and the cause is traffic share, not wording. Two fixes: run 50/50 so both arms get the same warmth, and report cache-hit rate per arm next to latency. Also keep per-call text, like the caller's name or today's date, at the end of the prompt. Put it at the top and no two calls share a prefix in either arm.

Sample-size math at mid-size volumes

This is where most prompt A/B tests fail: they stop at 200 calls. Here's the math for a two-sided test of two proportions at alpha 0.05 and 80% power.

n per arm = [ z(α/2)·√(2·p̄·(1−p̄)) + z(β)·√(p1(1−p1) + p2(1−p2)) ]² / (p2 − p1)²
z(α/2) = 1.960, z(β) = 0.842

Worked example. Baseline task success p1 = 0.80. You want to detect a lift to p2 = 0.83, so p̄ = 0.815:

  • First term: 1.960 × √(2 × 0.815 × 0.185) = 1.960 × 0.5491 = 1.0762
  • Second term: 0.842 × √(0.16 + 0.1411) = 0.842 × 0.5487 = 0.4618
  • n = (1.538)² / 0.0009 ≈ 2,629 calls per arm

That assumes every call is independent. Callers who call back aren't. The design effect is 1 + (m − 1) × ICC, where m is the average number of calls per caller during the test and ICC is how alike one caller's outcomes are. As an illustrative assumption, take m = 1.5 and ICC = 0.3. The design effect is 1.15, so you need about 3,024 calls per arm, or 6,048 total.

Now convert that to calendar time. Assume an illustrative mid-size line with 400 eligible calls a day:

Effect to detectSplitCalls needed (with 1.15 design effect)Days at 400 calls/day
80% to 85%50/502,0845.2, run 7
80% to 83%50/506,04815.1, run 21
80% to 85%90/105,78914.5, run 21
80% to 83%90/1016,80042

Two lessons. First, a three-point lift takes three weeks on a 400-call line. A one-point lift would take months, so don't run an online A/B for it; use the offline suite and accept the risk. Second, a cautious 90/10 split costs you a lot. With unequal arms, total sample scales by (1 + r)² / 4r, where r is the ratio between arms. At r = 9, that's 2.78 times the 50/50 total. Use a small split for safety during the first day or two, then move to 50/50 for the measurement. Only analyze data collected at the final split.

Round run time up to whole weeks. Monday calls aren't Saturday calls, and a test that covers five weekdays and no weekend measures a population you don't serve.

Sample size and run time for a voice agent prompt A/B test at 400 calls per day: 50/50 versus 90/10 splits for three and five point lifts in task success

Guardrails need their own power check. Detecting policy violations doubling from 2% to 4% takes about 1,141 calls per arm by the same formula, so a 3,024-per-arm test covers it. Detecting a rise from 2% to 2.5% would need far more. For rare, severe failures, use a hard threshold ("any confirmed PHI disclosure stops the test") rather than a significance test.

Peeking, interference, and novelty

Don't stop when the dashboard turns green

If you check a fixed-sample test every day and stop the first time p < 0.05, your real false-positive rate climbs well above 5%. Johari, Pekelis and Walsh show that p-values are unreliable under continuous monitoring and propose always-valid p-values that stay correct no matter when you look. Pick one: fix the sample size and look once, or use a sequential method built for peeking. Our Deepgram voice A/B guide walks through mSPRT with code. Guardrail checks for safety are different. Watch those daily and stop for harm, never for a win.

Interference: when arms share infrastructure

Classic A/B math assumes one arm can't affect the other. On a voice stack, they share a lot:

  • Rate limits. Both arms usually draw on one LLM tokens-per-minute quota. If B's prompt is 40% longer, it eats more quota, and at peak hours both arms get throttled. A's latency gets worse because of B, so the difference between arms understates the real cost of shipping B to everyone.
  • Prompt caches. Covered above. Unequal splits mean unequal cache warmth.
  • Shared queues. If B escalates more, the human queue gets longer for A's callers too. A's abandonment rises, and B looks relatively better.
  • Shared inventory. In scheduling, both arms book from the same slots. If B books more, A's callers find fewer slots. The test overstates B's lift.

You can't remove all of this at mid-size scale. Do three things: run 50/50, watch the shared resource (429 errors, queue wait, open slots) as its own metric, and treat lifts under two points with suspicion when inventory is shared.

Novelty and primacy

Most callers of a support line don't notice a prompt change, so novelty effects are smaller than in UI tests. Repeat callers are the exception. A reminder or collections line where people call weekly will see a change in tone. Some respond to the change itself, which fades. Split the results by first-time vs repeat callers and compare week one with week two. If the lift exists only in week one among repeat callers, it's novelty.

Blind vs deterministic sampling for human review

People search for this exact protocol, and the two terms mean different things. You want both.

Deterministic sampling is about which calls get reviewed. The selection rule is fixed before anyone sees results and is reproducible: for example, review every call where `sha256(call_id) mod 100 < 4`, stratified so each arm gets the same count. Anyone can rerun the rule and get the same list. The alternative, "pick some interesting calls," lets reviewers choose calls that confirm what they expect. Stratify by outcome too. If you only review failed calls, you learn how B fails, not whether B fails less.

Blind review is about what the reviewer sees. Strip the arm label, variant hash, and any prompt-specific greeting text that gives the variant away, then shuffle A and B calls into one queue. A reviewer who knows which call is "the new prompt" scores it differently, on purpose or not.

Put together, the protocol is:

1. Fix the sampling rule and per-arm quota before launch.

2. Draw the sample with the hash rule, stratified by arm and by outcome (success, escalation, failure).

3. Remove arm labels and give-away phrases. Shuffle into one queue.

4. Use a rubric of yes/no questions, not a 1 to 10 score.

5. Double-score 20% of calls and report agreement (Cohen's kappa) before trusting any difference.

6. Unblind only after all scoring is locked.

For sizing the review sample, see sampling live calls. For reviewer workflow, see human-in-the-loop evaluation.

There's a second meaning in offline work. "Deterministic sampling" can mean decoding at temperature 0 so each run repeats. As covered above, don't do that for an A/B. Sample at production settings and run k times.

Analyzing the live test

Analyze at the caller level. The simple two-proportion z-test below is fine when nearly every caller calls once. When repeats are common, the caller-clustered bootstrap gives an honest interval.

# online_ab.py (illustrative): two-proportion test + caller-clustered bootstrap
import math, random
from collections import defaultdict

def two_prop_z(s_a: int, n_a: int, s_b: int, n_b: int) -> tuple[float, float]:
    p_a, p_b = s_a / n_a, s_b / n_b
    p = (s_a + s_b) / (n_a + n_b)
    se = math.sqrt(p * (1 - p) * (1 / n_a + 1 / n_b))
    z = (p_b - p_a) / se
    p_value = math.erfc(abs(z) / math.sqrt(2))  # two-sided
    return p_b - p_a, p_value

def clustered_bootstrap(calls: list[dict], iters: int = 2000, seed: int = 7):
    """calls: [{"caller": "...", "arm": "A"|"B", "success": 0|1}, ...]"""
    by_caller = defaultdict(list)
    for c in calls:
        by_caller[(c["arm"], c["caller"])].append(c["success"])
    arms = {a: [v for (arm, _), v in by_caller.items() if arm == a] for a in "AB"}
    rng, diffs = random.Random(seed), []
    for _ in range(iters):
        rate = {}
        for a, clusters in arms.items():
            draw = [rng.choice(clusters) for _ in clusters]   # resample callers
            rate[a] = sum(map(sum, draw)) / sum(map(len, draw))
        diffs.append(rate["B"] - rate["A"])
    diffs.sort()
    return diffs[int(0.025 * iters)], diffs[int(0.975 * iters)]

Decision rules

Write these down before launch so the result decides, not the mood in the room.

Primary metricGuardrailsDecision
B better, interval excludes 0All within limitsShip B. Ramp 50% to 100% over a day, keep A warm for rollback
B betterAny guardrail breachedDon't ship. Fix the cause, rerun offline, then online
No significant differenceAll within limits, B cheaper or simplerShip B only if you set a non-inferiority margin up front and B is inside it
No significant differenceAny breachedKeep A
B worseAnyKeep A. Add the losing scenarios to the offline suite
Sample ratio mismatchAnyResult invalid. Fix assignment and restart

The last step matters most for the long run. Every live loss should become an offline scenario. Over a few experiments, the offline gate starts catching what used to need live traffic to find.

How to A/B test voice agent prompts, step by step

1. Write the experiment record. Variant bundle and hash, hypothesis, primary metric, guardrails with limits, minimum detectable effect, sample size, run length in whole weeks, and decision rules.

2. Run the offline paired A/B. Same scenarios, k = 5 runs per variant at production temperature, deterministic assertions first, position-swapped blind judge for the rest. Analyze paired differences and pass^k.

3. Gate. Block on any regression in compliance or payment scenarios, a worse pass^k, or a judge tie rate high enough to make the verdict meaningless.

4. Wire assignment. Salted hash of a stable caller key, chosen once at session start, logged as an exposure record. LiveKit: in the entrypoint or dispatch metadata. Pipecat: in the webhook router as a Stream parameter.

5. Start small, then measure at 50/50. A short low-percentage safety phase, then the planned split. Analyze only the 50/50 period.

6. Check the ratio daily, guardrails daily, primary once. Stop early for harm, never for a win, unless you use a sequential method.

7. Run blind, deterministic human review on a pre-set stratified sample.

8. Decide by the table, then backfill. Ship or not, and turn every live loss into an offline scenario.

Steps 2, 3, and 7 are where a team without an eval function usually runs out of time. Evalgent runs the offline paired suite with repeated runs and order-swapped scoring, and scores the live sample blind against your rubric, so the decision rests on evidence instead of ten test calls.

Frequently asked questions

How do you A/B test voice agent prompts?

Run both prompts offline on the same scenarios, k times each at production temperature, and compare paired differences and pass^k. If the new prompt passes, split live traffic by caller with a salted hash, pin each call to one arm at session start, and measure one primary metric plus guardrails over a sample size fixed in advance.

Should I randomize by call or by caller?

By caller. Call-level assignment lets one person hear both prompts, which is confusing and makes their calls dependent, so your effective sample is smaller than your call count. Hash a stable key, such as phone number or contact ID, with an experiment-specific salt. Avoid room names and call SIDs because they change every call.

How many calls do I need to A/B test a voice agent prompt?

To detect task success moving from 80% to 83% at 80% power and alpha 0.05, you need about 2,629 calls per arm. With repeat callers adding an illustrative 1.15 design effect, that's about 3,024 per arm. At 400 calls a day on a 50/50 split, plan three weeks.

Can an LLM judge pick the better prompt?

Only with controls. Zheng et al. (2023) found GPT-4 kept the same verdict after swapping answer order in just 65% of cases. Judge every pair in both orders, count a win only when both orders agree, hide variant labels, and report the tie rate. Prefer deterministic assertions wherever the outcome is checkable.

What is blind vs deterministic sampling in voice A/B evaluations?

Deterministic sampling fixes which calls get reviewed with a reproducible rule, such as a call-ID hash, set before results exist and stratified by arm and outcome. Blind review hides which variant produced each call and shuffles both arms into one queue. Use both together so neither selection nor scoring can favor the expected winner.

Why does my new prompt look slower in the A/B test?

Often it's the prompt cache, not the wording. A new prompt has a new prefix and starts cold, and a small traffic share keeps it cold more often. Run 50/50, report cache-hit rate per arm, and keep per-call details at the end of the prompt so the stable prefix can be shared.

Can I stop the A/B test early if B is clearly winning?

Not with a fixed-sample test. Checking daily and stopping at the first significant result inflates false positives. Either look once at the planned sample size or use a sequential method with always-valid p-values. Guardrails are different: watch them daily and stop immediately for harm.

Is offline A/B testing enough without live traffic?

For small expected effects or high-risk lines, often yes, because a one-point lift would take months to confirm live. Offline testing can't tell you how real traffic responds, though. Use the offline paired A/B as a gate for every change and save live tests for changes expected to move outcomes by several points.

The bottom line

A/B testing voice agent prompts works when the offline half pairs runs and swaps the judge's order, and the online half assigns by caller, pins every call, and sizes the sample before launch. Skip any of those and you'll ship prompts on noise, which on a phone line means real callers find the regression for you.

Related Articles