Open door for builders.
How to Benchmark a Decision Model on Your Own Call Data (Jev, OpenAI Decisions, LLM Judges)

On this page
Your voice agent makes dozens of small decisions on every call. Route to billing or support. Escalate or keep going. Allow the refund tool or block it. Until recently those went to a general-purpose LLM with a prompt and a JSON parser. Now there are models built only for this job.
Jev from TypeSafe shipped on September 15, 2026. OpenAI announced its Decisions API at DevDay on September 29, 2026. Both promise fast, structured answers instead of generated text. Both come with vendor numbers. None of those numbers were measured on your callers, your transcripts, your options or your network path.
This guide is a protocol you can run this week on Jev and any LLM judge, and on OpenAI Decisions the day it opens up: dataset, sample size, metrics with formulas, failure-mode tests and a pass or fail scorecard. For the product comparison, read OpenAI Decisions API vs Jev for voice agents.
What a decision model is, and what you are measuring
Decision model: a model that takes context (state) plus a question with a closed set of answers and returns one of those answers, ideally with a probability for each. It does not write prose.
Three kinds of systems compete for this job in a voice stack today.
| System | What it returns | Probabilities | Status (Oct 4, 2026) |
|---|---|---|---|
| Jev (TypeSafe) | Choice, Score or Noul answers; probabilities on all three, confidence on Choice and Score | Yes, documented | Public API; `POST https://api.typesafe.ai/v1/systemone` |
| OpenAI Decisions API | "Answers" to "user-defined questions with finite pre-defined answers" | Not documented | Limited preview; broad release "in the coming days" |
| LLM judge (any chat model) | Generated text you parse into a label | Only via token logprobs or a verbalized number | Generally available |
The OpenAI wording comes from the DevDay 2026 recap. As of October 4, 2026, OpenAI has not published an endpoint, schema, SDK method or pricing for Decisions, a gap Firecrawl's analysis of the docs and SDKs also found. So every test below runs against an interface, not a vendor.
Jev's response shape is documented. A Choice answer looks like this in the TypeSafe quickstart:
{
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
}That `probabilities` object is what makes calibration testing possible. A model that returns only a label can be scored for accuracy, but you can't check whether its confidence means anything or set a threshold that hands hard cases to a human.
Which decisions to benchmark
Benchmark each decision your agent makes, not "the model," because accuracy on one says little about another. Five decision families show up in almost every in-house voice agent.
| Decision family | Example question | Options | When it runs | Cost of an error |
|---|---|---|---|---|
| Routing | Which queue should handle this caller? | 3 to 12 queues | Mid-call, on the caller's turn | Wrong team, transfer loop |
| Escalation | Should a human take over now? | Yes / no / ask to confirm | Every caller turn | Angry caller kept with a bot, or wasted agent time |
| Tool gating | Is it safe to call `issue_refund` now? | Allow / block / confirm first | Before each tool call | Money moved wrongly |
| Compliance flag | Did the agent read the recording disclosure before collecting details? | Yes / no | Post-call | Regulatory exposure |
| Disposition | How did the call end? | Resolved / escalated / abandoned / voicemail | Post-call | Bad reporting, bad billing |
The first three run live inside a caller's turn, so latency matters. The last two run after the call, so cost dominates. For a deeper map of in-call decision points, see Jev for voice agent routing and escalation and Jev for tool calling.
Building the dataset from real calls
Cut the state at the decision point
A live routing decision is made on what the agent knew at that moment. If you hand the model the full call transcript, you leak the answer. The agent's next line ("Let me transfer you to billing") or the caller's later reply sits right there in the text.
For each sampled decision, store:
- `state`: the transcript prefix up to the end of the caller turn where the decision fired, plus only the structured fields your production code actually passes (account tier, open order count).
- `question` and `options`: the exact wording and option list used in production, stored as versioned text.
- `gold`: the adjudicated human label.
- `stratum` tags: decision type, call type, channel, noise flag, language, week.
For post-call decisions (disposition, compliance), the full transcript is the correct state. That's the only case where it is.
Use the transcript your agent saw
Runtime decisions are made on streaming STT output, errors included. Benchmark on a cleaner transcript and you overstate accuracy. Use the final transcript your agent logged at decision time as the primary state.
Then build a second copy of the state from an independent STT pass or a human-corrected transcript for 100 to 200 items. The accuracy gap between the two copies is your STT-induced error: it tells you whether a wrong route came from the decision model or a misheard word. Our guide on transcript vs audio evaluation covers when transcripts hide failures that only audio shows.
Human labelers should listen to the audio, not read the agent's transcript. Otherwise your gold labels inherit the same STT errors you are trying to measure.
Stratify, then oversample rare classes
Escalation might fire on 3% of turns. A random sample of 900 turns gives you about 27 positives. That's not enough to say anything about recall on the class that matters most.
Use stratified sampling:
1. Split by decision type. Each gets its own sample and its own score.
2. Within each type, set a floor per class, for example at least 100 examples of every option, by oversampling rare classes.
3. Spread across strata you care about: call type, phone channel (PSTN vs WebRTC), noisy vs clean, top accents or languages, and at least two different weeks.
4. Record the sampling weight of each item (population share divided by sample share).
Report two numbers per decision. Balanced metrics (macro-F1, per-class recall) use the stratified sample as is. Population metrics (accuracy as callers experience it) reweight each item by its sampling weight. A model can look strong on balanced metrics and still fail the 3% of calls that most need it, or the reverse.
Hold out a lock box
Split once: 70% development, 30% lock box. Tune question wording and option descriptions on the development set. Report final numbers only on the lock box, once per candidate.
The broader mechanics of building and refreshing a labeled set are in golden datasets for voice agent evaluation.
How many labeled examples you need
You need enough items that the confidence interval is narrower than the difference you care about. Start from the margin of error you can live with.
Sample size for a target margin
For one decision type, the number of items to estimate accuracy within ±E at 95% confidence is:
n = z^2 × p × (1 − p) / E^2 with z = 1.96Plug in an expected accuracy p. These are worked examples, not targets.
| Expected accuracy p | Margin ±E | Items needed |
|---|---|---|
| 0.90 | ±3 pts | 385 |
| 0.90 | ±2 pts | 865 |
| 0.95 | ±2 pts | 457 |
| 0.50 (worst case) | ±3 pts | 1,068 |
So 400 to 900 labeled items per decision type is the working range. Less than that and you can't tell a 92% model from a 95% one.
Report Wilson intervals, not plus-or-minus
The simple ±1.96 × √(p(1−p)/n) interval breaks near 0% and 100%, which is exactly where good decision models live. Use the Wilson score interval:
center = (p̂ + z²/2n) / (1 + z²/n)
half = z × √( p̂(1 − p̂)/n + z²/4n² ) / (1 + z²/n)
CI = center ± halfIllustrative: 870 correct out of 900 is 96.7%, Wilson 95% CI 95.3% to 97.7%. Now take the escalation class alone: 58 correct out of 60 is also 96.7%, but the Wilson CI is 88.6% to 99.1%. Same point estimate, very different evidence. Always print the interval next to the number.
Comparing two models on the same items
When you compare Jev to an LLM judge, both answer the same items, so use a paired test. Only items where they disagree carry information. McNemar's test works on those discordant pairs. The sample size to detect a true difference δ at α = 0.05 and 80% power is:
n = [ z_α/2 × √p_d + z_β × √(p_d − δ²) ]² / δ²Here p_d is the share of items where the two models disagree. Illustrative: if they disagree on 8% of items and you want to detect a 3-point accuracy difference, n ≈ 696 paired items. If they disagree on 10%, n ≈ 870. Pilot 100 items first to estimate p_d.
Double-labeling sets the human ceiling
A decision model can't be more accurate than your labels. If two trained people disagree on 8% of escalation decisions, a model scoring "92% vs one labeler" may be as good as a human.
Have two people label every item independently, from audio, using a written guide with boundary examples. Then measure agreement with Cohen's kappa:
κ = (p_o − p_e) / (1 − p_e)
p_o = observed agreement
p_e = Σ over labels of (share of rater A using label) × (share of rater B using label)Illustrative: two labelers agree on 92% of items, and chance agreement given their label mix is 55%. κ = (0.92 − 0.55) / 0.45 = 0.82.
What to do with the result:
- κ below about 0.6 on a decision. Your question is ambiguous. Fix the definition and the option descriptions before you benchmark any model. A model can't be measured against a target humans can't hit.
- Adjudicate disagreements. A third person resolves every split. The adjudicated label is gold.
- Use the human pair as the ceiling. Compute the model's κ against gold. If the model's κ is within a few points of the human-vs-human κ, more accuracy work on that decision has poor returns.
The labeling is the expensive part of the whole benchmark. Illustrative: 900 items × 2 labelers × 45 seconds of audio review each is about 22.5 labeler-hours per decision type. Model calls for the same run cost cents to a few dollars, as the cost section shows.

The seven tests
Each test answers one question a production owner asks. Run all seven on every candidate, using the same items, question wording and option text.
Test 1: Accuracy and macro-F1 per decision
Accuracy is the share of items where the predicted label equals gold. Report it per decision type with a Wilson interval.
Macro-F1 averages F1 across classes, so a rare class counts as much as a common one:
precision_c = TP_c / (TP_c + FP_c)
recall_c = TP_c / (TP_c + FN_c)
F1_c = 2 × precision_c × recall_c / (precision_c + recall_c)
macro-F1 = mean of F1_c over classesFor safety-type decisions, report recall on the dangerous class on its own line. A tool gate that misses 1 in 20 "block" cases can still post 97% accuracy if blocks are rare.
Test 2: Calibration (ECE, reliability diagram, Brier)
A model is calibrated when its 0.9 answers are right about 90% of the time. You need this if you plan to act on confidence, for example auto-routing above a threshold and confirming below it. Guo et al. (ICML 2017) showed that modern neural networks are often poorly calibrated out of the box, and that a one-parameter fix, temperature scaling, often repairs it. The same paper popularized the metric most teams use.
Expected Calibration Error. Sort items into M equal-width bins by top-label confidence (10 bins is common). For each bin, compare accuracy with mean confidence:
ECE = Σ over bins b of (|B_b| / n) × | acc(B_b) − conf(B_b) |Reliability diagram. Plot acc(B_b) against conf(B_b) per bin. Points below the diagonal mean overconfidence. Print the bin counts, since a bin with 6 items is noise.
Brier score. For multiclass decisions, Brier = mean over items of Σ_k (p_k − y_k)², where y is one-hot gold. Lower is better. Unlike ECE, it punishes both wrong answers and mushy probabilities, so it's a good single number for comparing two calibrated models.
Which "confidence" to bin? For Jev, the confidence page documents that `confidence` is derived from `probabilities`. For Choice it is (p_max − 1/n) / (1 − 1/n). That's a rescaling of the top probability, not a probability itself. For ECE, bin on p_max (the top probability). Use `confidence` for gating if you like, but don't treat a 0.78 confidence as "78% likely correct."
Where calibration is possible today:
| Candidate | Source of a probability | Weakness to test for |
|---|---|---|
| Jev | Documented `probabilities` per option | Shift under option reorder (Test 5) |
| OpenAI Decisions | Unknown; not in public material | If no probabilities, only accuracy and consistency are measurable |
| LLM judge, logprobs | Token log probabilities on a one-token label | Only works if your model exposes logprobs; multi-token labels need care; reasoning-style models may not expose them |
| LLM judge, verbalized | "Give a confidence 0 to 100" | Overconfidence and coarse values (multiples of 5 or 10) |
| LLM judge, vote | Share of k samples that agree | Costs k calls; collapses if the model is highly consistent |
The research on LLM confidence is mixed. Tian et al. (EMNLP 2023) found that for RLHF-tuned chat models, verbalized confidences were often better calibrated than token probabilities, cutting ECE by a relative 50% on their QA benchmarks. Xiong et al. (ICLR 2024) found LLMs tend to be overconfident when they verbalize confidence, and that sampling consistency helps. Neither tested call-center routing. Treat both as hypotheses your benchmark settles.
Test 3: Coverage at a confidence threshold
This is the test that turns a benchmark into an automation plan. For a threshold τ, auto-decide only items with confidence ≥ τ and send the rest to a fallback (ask the caller to confirm, route to a human, or call a slower model).
coverage(τ) = share of items with conf ≥ τ
selective_accuracy(τ) = accuracy on those items onlySweep τ from 0.5 to 0.99 and plot both. Then read off the operating point: "At τ = 0.9, the model auto-decides 81% of routing turns at 98.6% accuracy; the other 19% go to a confirmation prompt." (Illustrative numbers.) Two models with the same raw accuracy can differ a lot here. The one with better-ordered confidence automates more calls at the same error rate.
Without probabilities there is no curve. That is a scorecard item, not a footnote. If OpenAI Decisions launches without per-answer probabilities, you can still benchmark accuracy, consistency and latency. You lose selective automation unless you build a confidence proxy, such as agreement with a second model.
Test 4: Run-to-run consistency
Send every item k times (k = 5 is enough to spot trouble) with identical inputs. Measure:
flip_rate = share of items where the k labels are not all identicalFor Score-type questions, also report the per-item variance of the score across runs, averaged over items.
The reason this matters: a flip on a live decision means two callers with the same words get different treatment, and a flip on a post-call judge means your dashboard moves when nothing changed. LangChain's Jev-as-a-judge study ran each judge 100 times per case. Jev's mean per-case quality-score variance was 0.0000149. GPT-5.6 Luna's was 433 times higher, GPT-5.6 Terra's 913 times higher and Claude Sonnet 4.6's 92 times higher. On the binary pass/fail decision, Jev matched the human oracle on all 500 repeated decisions, Terra on 99.8%, Luna on 96.4% and Claude on 80.0%.
Read the limits too: five weather-agent cases, one human labeler and provider-default sampling settings for the LLM judges. The authors call it narrow and early. It shows the size of gap that's possible, not the gap on your routing questions. The longer discussion of judge drift is in the limits of LLM-as-judge.
Low variance is not accuracy. A model that is consistently wrong on a class will post a zero flip rate. Read Test 4 next to Test 1.
Test 5: Option-order sensitivity
Reorder the answer options and see if the answer changes. The evidence for LLMs is strong. Pezeshkpour and Hruschka (2023) measured performance gaps of roughly 13% to 75% across benchmarks when answer options were reordered, and traced the effect to cases where the model is torn between its top two or three choices. Zheng et al. (ICLR 2024) tested 20 LLMs and found a "selection bias" toward specific option IDs such as "A," rooted in token-level priors. They proposed estimating that prior by permuting options on a small sample.
Decision models aren't exempt. On Hacker News, a commenter argued that Jev "can output drastically different probabilities if you simply reorder the list of choices," and that its confidence is "just a formula applied to probabilities." The second point matches TypeSafe's own documentation. On the first, TypeSafe's jev-1.13 jaggedness page says that "in some cases" option order can affect the answer and the model "leans toward the option that comes first," and tells users to reorder options to check consistency. Run that check yourself.
Method:
1. For each item, build m variants of the option list: original, reversed, and m − 2 random shuffles (m = 4 is a good start). Keep option keys and descriptions identical; only order changes.
2. Map every response back to canonical option order.
3. Compute:
order_flip_rate = share of items whose top label differs across the m variants
max_prob_shift = mean over items of max over options of (max_v p_v(option) − min_v p_v(option))
first_slot_rate = share of answers that pick whichever option was listed firstCompare first_slot_rate with the share of gold labels that happened to be first. A gap shows position bias.
If order flips exceed your threshold, rewrite the near-tie options so they're less confusable (the root-cause fix), or average probabilities across two or three permutations sent in parallel. Freezing one canonical order hides the bias rather than removing it.
Test 6: Irrelevant-context robustness
Real decision state is messy. Earlier turns about something else, a CRM blob, a long account history. Shi et al. (ICML 2023) built GSM-IC, math problems padded with irrelevant sentences, and found model accuracy "dramatically decreased when irrelevant information is included." TypeSafe's jaggedness page says the same about Jev: accuracy "falls as the state grows with content unrelated to the decision," and recommends filtering state in code first.
Method: for each item, build three state sizes.
- Minimal: only the turns and fields the decision needs.
- Production: what your code sends today.
- Padded: production state plus distractors drawn from other real calls (unrelated earlier turns, unrelated account notes) up to 2 to 4 times the production size.
Report accuracy at each size and the drop:
distractor_drop = acc(minimal) − acc(padded)A large drop between minimal and production is a free win: trim your state. A large drop between production and padded is a warning about long calls, where state naturally grows. Check context limits too. Jev is listed at 32K tokens on OpenRouter, and TypeSafe's docs list 64K; long calls with full history can approach those.
Test 7: Latency and cost, measured by you
Vendor latency numbers measure something real, but not your path. OpenAI's DevDay slide showed 150 ms for Decisions against 1.6 s for GPT-6 Luna, conditions unstated. TypeSafe quotes 70 to 500 ms for Jev. OpenRouter's telemetry for Jev showed P50 0.21 s, P95 0.34 s and P99 0.75 s as of October 1, 2026. None of them is your agent server, in your region, sending your state size.
Measure client-side, from the same region and host type that runs your agent:
- Clock: wall time from just before the HTTP request to the parsed answer, including TLS and JSON parsing.
- Warm vs cold: report warm numbers after the first 20 calls, and cold-start (first call on a new connection) separately. The first caller after a deploy pays the cold path.
- Concurrency: 40 simultaneous calls each making a decision every 6 seconds is about 7 decisions per second, with bursts. Run 1×, 3× and 10× that rate and count 429s and timeouts. See concurrency failures in voice agents.
- Volume: at least 2,000 timed calls per condition. Below about 1,000, p99 is decided by a handful of requests.
- State size: use production-size state from Test 6, not a one-line test string.
Report p50, p95, p99, error rate and timeout rate, ideally across a full day.
Cost per 1,000 decisions. Illustrative, assuming 1,500 input tokens of state and question per decision:
| Candidate | Price basis | Cost per 1,000 decisions |
|---|---|---|
| Jev | $0.042 per 1M input tokens, output free | 1.5M × $0.042/1M = $0.063 |
| LLM judge on GPT-6 Luna (yardstick) | $0.10 per 1M input, $0.50 per 1M output; assume 150 output tokens | $0.15 + $0.075 = $0.225 |
| OpenAI Decisions | Not published | Unknown; do not assume Luna rates |
Add retries and any permutation or vote calls you plan to make in production. The benchmark run itself is cheap: 900 items × 12 calls (1 base, 4 repeats, 4 permutations, 3 state sizes, with overlap) is about 10,800 calls. At 1,500 tokens each, that's roughly 16M tokens, which comes to about $0.68 on Jev at list price.
The latency budget: where a decision fits in a turn
Whether a decision API's p95 is good enough depends on where it sits in the turn. Here's the arithmetic for a cascaded agent, with illustrative hop values you should replace with your logged numbers.
| Hop | Assumed p95 |
|---|---|
| End-of-turn detection after the caller stops | 300 ms |
| STT finalization | 100 ms |
| LLM time to first token | 400 ms |
| TTS time to first audio byte | 150 ms |
| Network and telephony legs | 100 ms |
| Total without a decision call | 1,050 ms |
If your target is about 1,100 ms to first audio, a decision call placed in series before the LLM has 50 ms of budget. No network API hits that at p95. Even OpenAI's 150 ms headline figure would push the turn over.
Placed in parallel, the picture changes. Fire the decision and the LLM request together on the final transcript. The decision is free if it returns before you need it, for example before the LLM emits a tool call. Its budget becomes the LLM's time to first token: about 400 ms here. Jev's OpenRouter P95 of 340 ms would fit; its P99 of 750 ms would not, so somewhere between 1 and 5 turns in 100 would wait.
Two non-obvious pieces of tail math:
- Several decisions per turn. If a turn fires three independent decisions in parallel and each has p95 = X, the chance all three finish within X is 0.95³ ≈ 0.857. Your turn-level p95 is really a p86. To keep the turn at p95, each call needs to meet the budget at about its p98.3 (0.95^(1/3) ≈ 0.983). Read p99 numbers, not p95, when decisions fan out.
- Many turns per call. On a 10-turn call, the chance at least one turn hits a p99 tail is 1 − 0.99¹⁰ ≈ 9.6%. With a p95 tail, it's 1 − 0.95¹⁰ ≈ 40%. Callers judge the call by its worst pause.
Set a hard timeout at the budget, with a safe default action (for example, "ask the caller to confirm") when the decision is late, and count timeouts as benchmark failures.

A provider-agnostic harness in Python
Three layers: an interface, one adapter per candidate, and vendor-blind metrics. The code is simplified; add retries, rate limiting and logging before running at volume.
The interface
# harness/providers.py (simplified)
import time
from dataclasses import dataclass
from typing import Optional, Protocol
@dataclass
class DecisionResult:
label: str
probs: Optional[dict[str, float]] # None when the provider gives no probabilities
latency_ms: float
class DecisionProvider(Protocol):
name: str
def decide(self, state: str, instructions: str,
options: dict[str, str]) -> DecisionResult:
"""options maps option key -> description, in the order to present."""
...Jev adapter
This uses the documented `typesafe-sdk` (`pip install typesafe-sdk`, Python 3.10+). The client reads `TYPESAFE_API_KEY` from the environment. Pin the model version for a benchmark so results don't move when `jev-latest` changes.
from typesafe_sdk import Choice, TypeSafeClient
class JevProvider:
name = "jev"
def __init__(self, model: str = "jev-1.13"):
self.client = TypeSafeClient(model=model)
def decide(self, state, instructions, options):
t0 = time.perf_counter()
resp = self.client.system_one(
state=state,
questions={"d": Choice(instructions=instructions, criteria=options)},
)
latency = (time.perf_counter() - t0) * 1000
ans = resp.answers["d"]
return DecisionResult(label=ans.choice,
probs=dict(ans.probabilities),
latency_ms=latency)LLM judge adapter with logprobs
This uses the OpenAI Chat Completions `logprobs` and `top_logprobs` parameters (top_logprobs is capped at 20). It only works with a model that returns logprobs; check your model's docs. Options are numbered so each label is a single token, which is also why Test 5 matters here: numbers are positions.
import math
from openai import OpenAI
class LLMJudgeProvider:
name = "llm_judge"
def __init__(self, model: str):
self.client = OpenAI()
self.model = model # a model that supports logprobs
def decide(self, state, instructions, options):
keys = list(options)
menu = "\n".join(f"{i+1}. {options[k]}" for i, k in enumerate(keys))
prompt = (f"{instructions}\n\nOptions:\n{menu}\n\n"
f"Context:\n{state}\n\nAnswer with the option number only.")
t0 = time.perf_counter()
resp = self.client.chat.completions.create(
model=self.model,
messages=[{"role": "user", "content": prompt}],
max_tokens=1, logprobs=True, top_logprobs=20, temperature=0,
)
latency = (time.perf_counter() - t0) * 1000
first = resp.choices[0].logprobs.content[0].top_logprobs
raw = {k: 0.0 for k in keys}
for t in first:
tok = t.token.strip()
if tok.isdigit() and 1 <= int(tok) <= len(keys):
raw[keys[int(tok) - 1]] += math.exp(t.logprob)
total = sum(raw.values()) or 1.0
probs = {k: v / total for k, v in raw.items()}
return DecisionResult(label=max(probs, key=probs.get),
probs=probs, latency_ms=latency)Renormalizing over valid option tokens hides how much mass went to junk tokens. Log the pre-normalization total too; if it's often well below 1, the model is not following the format and your probabilities are less trustworthy.
OpenAI Decisions placeholder
class OpenAIDecisionsProvider:
"""PLACEHOLDER. As of Oct 4, 2026 OpenAI has not published the Decisions API
endpoint, request/response schema, SDK method or whether probabilities are
returned. Fill this in from the official docs when they ship. Do not guess."""
name = "openai_decisions"
def decide(self, state, instructions, options):
raise NotImplementedError("Fill in when OpenAI publishes the Decisions API contract.")If the launched API returns no probabilities, set `probs=None`. The metric code below skips calibration and coverage for that provider and reports them as "not measurable," which is the honest result.
Metrics with numpy
# harness/metrics.py (simplified)
import numpy as np
Z = 1.959964
def wilson_ci(correct, n, z=Z):
p = correct / n
d = 1 + z**2 / n
c = (p + z**2 / (2 * n)) / d
h = z * np.sqrt(p * (1 - p) / n + z**2 / (4 * n**2)) / d
return c - h, c + h
def cohens_kappa(a, b):
a, b = np.asarray(a), np.asarray(b)
p_o = np.mean(a == b)
p_e = sum(np.mean(a == l) * np.mean(b == l) for l in np.union1d(a, b))
return (p_o - p_e) / (1 - p_e)
def ece(conf, correct, n_bins=10):
edges = np.linspace(0, 1, n_bins + 1)
idx = np.clip(np.digitize(conf, edges[1:-1]), 0, n_bins - 1)
return float(sum((idx == b).mean() * abs(correct[idx == b].mean() - conf[idx == b].mean())
for b in range(n_bins) if (idx == b).any()))
def brier(prob_matrix, gold_idx):
onehot = np.eye(prob_matrix.shape[1])[gold_idx]
return float(np.mean(np.sum((prob_matrix - onehot) ** 2, axis=1)))
def coverage_curve(conf, correct, thresholds=np.arange(0.5, 1.0, 0.05)):
return [(float(t), float((conf >= t).mean()),
float(correct[conf >= t].mean()) if (conf >= t).any() else float("nan"))
for t in thresholds]
def flip_rate(label_runs): # shape (n_items, k)
return float(np.mean([len(set(r)) > 1 for r in label_runs]))
def max_prob_shift(prob_runs): # shape (n_items, m_variants, n_options), canonical order
return float(np.mean((prob_runs.max(axis=1) - prob_runs.min(axis=1)).max(axis=1)))
def latency_pcts(ms):
return {f"p{q}": float(np.percentile(ms, q)) for q in (50, 95, 99)}The run loop
import random
def run_candidate(provider, items, k_repeats=5, m_perms=4):
out = []
for it in items: # it: state, instructions, options, gold
keys = list(it["options"])
base = [provider.decide(it["state"], it["instructions"], it["options"])
for _ in range(k_repeats)]
perms = []
for v in range(m_perms):
order = keys[::-1] if v == 1 else random.sample(keys, len(keys)) if v else keys
r = provider.decide(it["state"], it["instructions"],
{k: it["options"][k] for k in order})
perms.append(r)
out.append({"id": it["id"], "gold": it["gold"], "base": base, "perms": perms})
return outRun distractor variants as separate item sets so their latency doesn't pollute base numbers. Store raw responses as JSON lines with model version, question version and timestamp, so you can re-score later without new calls.
The scorecard: pass thresholds for live calls
The thresholds below are a starting template. They are illustrative, not industry standards. Tighten them for money-moving and compliance decisions; loosen them for post-call analytics.
| # | Test | Metric | Live in-call decision (routing, escalation, tool gate) | Post-call decision (disposition, compliance) |
|---|---|---|---|---|
| 1 | Accuracy | Lower bound of Wilson 95% CI | ≥ your target (e.g., 0.93) | ≥ your target (e.g., 0.95) |
| 1 | Class balance | Macro-F1; recall on dangerous class | Macro-F1 ≥ 0.85; dangerous-class recall ≥ 0.97 | Compliance-miss recall ≥ 0.98 |
| 1 | Human ceiling | κ(model, gold) vs κ(human, human) | Within 0.05 | Within 0.05 |
| 2 | Calibration | ECE on p_max (10 bins); Brier | ECE ≤ 0.05 | ECE ≤ 0.05 |
| 3 | Selective automation | Coverage at 98% selective accuracy | ≥ 70% of turns | ≥ 85% of calls |
| 4 | Consistency | Repeat flip rate, k = 5 | ≤ 1% | ≤ 1% |
| 5 | Order robustness | Order-flip rate, m = 4 | ≤ 2% | ≤ 2% |
| 6 | Distractors | distractor_drop at production size | ≤ 2 pts | ≤ 3 pts |
| 7 | Latency | Client-side, in your region, at peak concurrency | p99 ≤ parallel budget (e.g., ≤ 400 ms); timeouts ≤ 0.1% | p95 ≤ 2 s |
| 7 | Cost | Per 1,000 decisions incl. retries and permutations | Within budget | Within budget |
A candidate that returns no probabilities gets "not measurable" on rows 2 and 3. Decide up front whether that's disqualifying for each decision. For a live tool gate, it usually is, because you can't build a confirm-when-unsure path without a confidence signal.

Reading the results and choosing a model
Accuracy CIs overlap. If the Wilson intervals overlap and McNemar's test on discordant pairs isn't significant, the models are tied on accuracy at your sample size. Decide on calibration, latency and cost, not on a 0.4-point gap.
One model is more accurate, the other is calibrated. Compare coverage at your target selective accuracy, not raw accuracy. A calibrated model at 94% accuracy that can auto-decide 85% of turns at 99% is often more useful than an uncalibrated 96% model you can't threshold.
Low flip rate, mediocre accuracy. The model is stably wrong in the same places. Check the confusion matrix; often two options overlap in meaning. Rewrite them on the development split, then confirm on the lock box.
High order-flip rate. The model sees two options as near-ties. Fix the options first; if flips persist, average across permutations or switch models for that decision.
Large distractor drop. Trim state in code before you change models. TypeSafe's guidance, "filter first; send only what the question needs," applies to LLM judges too.
Passes at p95, fails at p99. Move the decision off the critical path if you can. If not, the model fails for that live decision even if it's the most accurate.
You can mix models by decision. A fast, calibrated decision model on live routing and tool gating, with an LLM judge for open-ended post-call review, is a common split. The trade-offs are covered in Jev vs LLMs for voice agents and Jev as a judge.
Pitfalls that invalidate a decision API benchmark
Label leakage. Full-call transcripts for in-call decisions; labelers who read the agent's own routing line; gold labels taken from the agent's logged decision. Each one makes the test easier than production. Cut state at the decision point and label from audio.
Prompt and option wording drift. You tune the question text for one model during development, then compare on that wording. The tuned model wins. Freeze one question version per decision before the lock-box run and use it for every candidate. If you must tune per model, give each the same number of tuning rounds on the development split and report that.
Comparing a vendor latency claim with your measured p50. A 150 ms slide figure, a 70 to 500 ms range and an OpenRouter P50 of 0.21 s are measured from different places, on different inputs. Compare only numbers you measured, with the same client, region, state size and concurrency.
Not testing under concurrency. Single-threaded benchmarks show the best case. Shared APIs rate-limit, queue and stretch tails under load. Run at 1×, 3× and 10× your peak decision rate, and count 429s and timeouts as failures.
Scoring the model on decisions code should make. TypeSafe's jaggedness page says Jev "is not a calculator" and recommends keeping arithmetic and date comparison in code. Move those items to code before you benchmark.
Mistaking consistency for correctness. A highly consistent model with a systematic error repeats it on every call. Low flip rate is necessary for a live decision, not sufficient.
Not re-running after a version change. `jev-latest` and model aliases move, and call mix shifts by week and season. Pin versions, record them with results, and re-run when the provider ships a new version. For ongoing scoring in production, see why you should score every call.
How to run the benchmark in one week
1. Day 1: define decisions. Pick two to four decision types, copy the production question and option text, and write a labeling guide with two boundary examples per option.
2. Day 1: pull and cut the sample. Stratify with a 100-per-class floor, cut state at the decision point, and keep audio links for labelers.
3. Days 2 to 3: double-label from audio. Staff enough labelers for the hours. Compute κ, rewrite any decision below about 0.6, and adjudicate splits to get gold.
4. Day 3: split and freeze. 70% development, 30% lock box. Freeze question version 1.
5. Day 4: wire adapters. Jev with a pinned model version, your LLM judge with logprobs or vote confidence, and the Decisions placeholder. Smoke-test 20 items each.
6. Day 4: run Tests 1 to 6 on the development split. k = 5 repeats, m = 4 permutations, three state sizes. Give every candidate the same tuning rounds.
7. Day 5: run latency and load from your agent's region: 2,000+ timed calls per candidate, warm and cold, at 1×, 3× and 10× peak.
8. Day 5: run the lock box once and compute every scorecard row, with Wilson intervals and McNemar tests.
9. Day 5: choose per decision. Set the confidence threshold from the coverage curve, the timeout and the fallback action.
10. After launch: re-run after any model, STT, prompt or option change, and add fresh calls quarterly.
When OpenAI publishes the Decisions contract, fill in the adapter and re-run steps 5 through 9. Nothing else changes.
How independent evaluation fits
You can run all of this in-house. The hard parts aren't the code. They're labeling from audio at volume, keeping the lock box sealed while engineers tune prompts, and producing a result a vendor or compliance reviewer accepts as neutral.
Evalgent, an independent evaluation platform for voice agents, does that work: it builds the double-labeled set from your recordings, runs every candidate through the same seven tests and load profile, and re-runs when a vendor ships a new version. It doesn't sell any of the models it scores. The wider method is in benchmarking voice agents on your own data.
Jev can't abstain: always include an escape option
Jev always returns one of the options you define. It has no built-in "I don't know." When a call transcript holds no evidence for any option, such as a two-word turn, a dropped call or a caller talking about something else, Jev still picks the least-wrong option, and that label looks just like a confident answer. TypeSafe's own Choice docs recommend adding an "other" or "none of the above" option when the list may not cover every input.
Build this into the benchmark. Include gold examples where the correct answer is the escape option, and measure how often each provider picks it when it should and avoids it when it shouldn't.
Four rules keep this from hurting you on live calls:
- Add an explicit escape option such as `not_enough_information`, `not_applicable` or `none_of_the_above` to every Choice question where evidence can be missing. Describe it as concretely as the real options, because Jev reads only the descriptions, not your option keys.
- Ask whether the evidence exists first. A Noul like "does the caller state a reason for calling?" before the Choice question separates "no evidence" from "evidence for option B".
- Check the distribution, not just the label. In one independent edge-case test, a forced choice with no valid option still returned a pick at 0.52 probability but only 0.04 confidence (jev-exploration). The numbers flagged the problem; the label did not. Set your human-review cutoff from your own labeled calls, because an independent study on data unlike Jev's training found its confidence can run high (jev-ood-calibration).
- Track the escape-option rate. If it climbs, your transcripts, STT or option set have drifted away from what callers actually say.
Frequently asked questions
How many labeled examples do I need to benchmark a decision model?
Per decision type, about 385 items give ±3 points at 90% accuracy and 865 give ±2 points, at 95% confidence. Rare classes need their own floor, for example 100 examples each. To compare two models on the same items, plan around 500 to 900 paired items, depending on how often they disagree.
Can I measure calibration for OpenAI's Decisions API?
Not yet. As of October 4, 2026, OpenAI has not documented whether Decisions returns probabilities or confidence. If it returns only a label, you can measure accuracy, consistency, order sensitivity and latency, but not ECE, Brier or coverage curves. Fill in the adapter when official docs ship and re-check.
Is Jev's confidence the same as its probability?
No. TypeSafe documents that confidence is computed from the probabilities. For a Choice it rescales the top probability so an even split is 0 and certainty is 1. For calibration metrics, bin on the top probability. Use confidence, or your own statistic, for gating thresholds after you validate it.
How do I get calibrated confidence from an LLM judge?
Three options. Read token logprobs on a single-token label, if your model exposes them. Ask for a verbalized confidence. Or sample several times and use the vote share. Research disagrees on which is best, and verbalized scores often run overconfident, so measure ECE for each method on your own labeled items.
What is a good flip rate for a live routing decision?
As a starting threshold, under 1% of items should change label across five identical runs, and under 2% across four option orders. These are illustrative bars, not standards. Tighten them for tool gates that move money or data, because a flip there means two callers with the same request get different outcomes.
Should I test on the agent's transcript or a clean transcript?
Use the transcript your agent saw at decision time as the main test, because that is what the model gets in production. Add a clean or independent-STT copy for a subset. The accuracy gap between the two tells you how much error comes from speech recognition rather than the decision model.
Why measure latency myself when vendors publish numbers?
Vendor figures come from different setups. OpenAI's 150 ms slide figure has unstated conditions, TypeSafe quotes 70 to 500 ms, and OpenRouter telemetry shows a P95 of 0.34 s for Jev. Your agent pays the latency from its own region, with its own state size and concurrency. Only that number belongs in a turn budget.
Can one model handle every decision in my voice agent?
Sometimes, but benchmark per decision first. Results often split: a fast, calibrated decision model wins on live routing and tool gating, while an LLM judge handles open-ended post-call review. Pick the winner per decision, with its own threshold, timeout and fallback, rather than one model for everything.
The bottom line
A decision model benchmark that uses your own calls, double-labeled gold, Wilson intervals and seven tests tells you more than any vendor chart, and it costs far less in API calls than in labeling time. Build it around a provider-agnostic interface now, run it on Jev and your LLM judge this week, and drop in OpenAI Decisions the day its contract is public.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more