Open door for builders.
Jev Accuracy in Voice Agent Evals: Where It Fails on Calls

On this page
You wire Jev into your post-call pipeline. Every call gets a QA rubric: did the agent verify identity, read the recording disclosure, confirm the appointment, resolve the issue. Jev answers each question in a fraction of a second for a fraction of a cent. The dashboard fills up.
Then someone listens to a call that Jev failed on "confirmed the appointment." It lasted nine seconds. The caller said "wrong number" and hung up. There was no appointment to confirm. Jev did not make a typo or hallucinate a new label. It did what it is built to do: pick one of the options you gave it.
This guide answers one question for teams that run their own voice agent: can you trust Jev to score your calls? The answer is yes for a large share of decisions, provided you design around one structural limitation. We cover what the public accuracy evidence actually shows, where Jev breaks on real call transcripts, what forced answers do to your pass rates, the fixes, working Python against the documented SDK, and how to test all of it on your own calls.
What the evidence says about Jev accuracy
Jev is TypeSafe's first System One model. It takes a state (your transcript plus metadata) and typed questions, and returns answers with probabilities instead of generated text. There are three question types: Choice (pick one option), Score (a position on ordered levels) and Noul (the probability a yes/no statement is true). If that is new, start with Jev for voice agents.
Four kinds of evidence exist so far. They measure different things, so read them side by side.
| Source | Who ran it | Task and size | Headline result | What it cannot tell you |
|---|---|---|---|---|
| LangChain, Jev-as-a-Judge | LangChain team | 5 fixed weather-agent runs, 100 repeats each (500 decisions) | Pass/fail matched the human oracle on 500 of 500; GPT-5.6 Terra 99.8%, GPT-5.6 Luna 96.4%, Claude Sonnet 4.6 80.0%. Quality-score variance 92 to 913 times lower than the LLM judges. $0.34 total vs $28.17 for Claude; 0.44 s and $0.00035 per call | Only five unique cases, so agreement is about repeatability on a narrow task. LLM judges ran at provider defaults; the Jev version was not recorded |
| TypeSafe launch post | TypeSafe | Four internal "workflow evals" | 193.6x faster and 444.6x cheaper than frontier LLMs. TypeSafe itself writes: "we expect that these are on the higher end of real world gains" | Reference answers are the average of two frontier LLMs, not human ground truth (TypeSafe says so) |
| jev-exploration | One independent tester | 12 support tickets; 4 edge-case probes | Department routing 10 of 12 (11 of 12 after a label revision), urgency 12 of 12 | n=12 and n=4. The author says the 95% intervals overlap a trivial baseline (majority class got 5 of 12), so this is not accuracy evidence |
| jev-ood-calibration | One independent tester | 900 synthetic support tickets Jev cannot have seen, 3 question types | Queue 89.0%, "is the customer angry" 91.7%, priority (a hidden policy rule) 44.7% | Synthetic text, one task family. Tells you about calibration out of distribution, not about call transcripts |
Credit where it is due. Every study that measured it found Jev highly consistent. LangChain's variance gap is large. TypeSafe's own self-consistency cookbook ran an 8-question moderation rubric 15 times: raw label agreement was 90.8%, and with a rule that sends any answer under 0.60 top probability to a human, agreement rose to 99.2% with 74.2% of answers handled automatically. The cost is low enough that you can score every call instead of a sample.
Two caveats run through all of it. First, consistency is not correctness. A judge that is reliably wrong produces bad data at scale, and LangChain says exactly that in its write-up. Second, none of the studies above used call transcripts. Accuracy swings by task in the independent ledger that jev-exploration maintains, from 98.3% on spam to 62.6% on phishing. Your calls are a task nobody has published.
The core failure: Jev cannot say "I don't know"
This is the mechanism to understand before you trust any number Jev gives you.
A Choice answer is a probability distribution over the options you supplied. The probabilities sum to 1. The `choice` field is "the option with the highest probability," per the Choice docs. The model cannot return an option you did not offer. So if none of your options fit the call, Jev still returns one of them. It picks the least-wrong one.
That is the price of type safety. TypeSafe's launch post is candid about it: its 0% hallucination figure carries the footnote "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." A valid label is guaranteed. A true label is not. TypeSafe's System One page says the same about probabilities: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct."
What the edge-case probes showed
The jev-exploration author ran four probes built to have no good answer. These are single probes from one tester, and the raw responses were not kept, so treat them as illustrations of behavior, not rates:
| Probe | What was asked | Jev's output |
|---|---|---|
| Forced Choice, no valid option | Pick between two options that were both wrong | Winner at 0.52, confidence 0.04 |
| Unknowable Noul | "Will the user renew?" with no evidence either way | 0.49 |
| Contradictory evidence | State supports both yes and no | 0.40 |
| Question unrelated to state | Asked about something not in the state | Answered from world knowledge at 0.04 |
The first row is the one that bites call scoring. TypeSafe computes Choice confidence as (p_max − 1/n) / (1 − 1/n), per the confidence docs. With two options, 0.52 works out to (0.52 − 0.5) / 0.5 = 0.04. The distribution did flag the problem. The answer did not. If your pipeline stores only `choice`, your database says "pass" or "fail" with nothing to show it was a coin flip.
The fourth row is worse for Noul questions. A Noul has no third outcome. "Did the agent verify the caller's identity?" on a six-second call returns a value near 0, which looks exactly like a confident "no." It is literally true that no verification happened. It is wrong for a compliance metric, because verification was never due.

Why this is an old problem with a known fix
Machine reading had the same failure. The SQuAD 2.0 paper, "Know What You Don't Know", added over 50,000 unanswerable questions to a benchmark where every question used to have an answer. A strong system that scored 86% F1 on the original set scored 66% F1 once it also had to decide when to abstain. Evaluation sets without "no answer" cases overstate accuracy. Your QA rubric is the same: if your test calls all have an appointment, you never see how the "appointment confirmed" item behaves on calls without one.
The fix is equally old. Selective classification, or the reject option, lets a model decline when it is not sure, trading coverage for lower error. Geifman and El-Yaniv showed you can pick a target error rate and reject inputs until you hit it; on ImageNet they guaranteed 2% top-5 error at almost 60% coverage. Jev gives you the raw material (a full distribution) but not the reject decision. You build it two ways: an explicit escape option the model can choose, and a threshold in your code. TypeSafe recommends the first one directly: "Add an `other` or `none of the above` option when the list might not cover every input, so the model can say none of the others fit."
Jev confidence score calibration when calls look unfamiliar
The second half of the abstain problem is whether the probabilities warn you. The best public test of that is jev-ood-calibration, which generated 900 support tickets from a script so Jev could not have seen them. One of three questions, ticket priority, depended on a policy rule that is not in the text (customer tier bumps priority). No model can recover that from the ticket. It is the closest public analog to a rubric item that depends on something outside your transcript.
What it found, as published on September 19, 2026:
| Measure | Result | What it means |
|---|---|---|
| Accuracy on the unknowable priority rule | 44.7% (chance is 25%) | Common sense gets partway; the rule itself is invisible |
| Average stated probability on that question | 0.74 | Jev did not signal that it could not know |
| ECE across all 900 items | 0.107, against a noise floor of 0.024 (4.4x) | Real miscalibration, not sampling noise |
| Refit temperature, all 900 | 2.74 | Probabilities were overconfident overall |
| Refit temperature by type | Choice 3.29, Score 3.40, boolean 0.66 | Choice and Score overconfident; booleans underconfident |
| Exact zeros (OpenBookQA run) | 1,051 of 2,000 option probabilities were exactly 0.00 | One item put 0.00 on the correct answer; no temperature fix can repair a hard zero |
The temperature refit is the method from Guo et al., "On Calibration of Modern Neural Networks": find the single scalar that makes probabilities honest after the fact. T above 1 means the model is overconfident, below 1 means underconfident. The author's practical advice follows from the sign flip: calibrate per question, not per model, and do not threshold on the `confidence` field, which was never better than the top probability in that study and sometimes much worse.
jev-exploration adds two more findings. Every probability Jev returned in its 281 committed responses sat on a 0.01 grid. And on its own 800-item set, Noul probabilities were compressed toward the middle, so a 0.9 threshold gave a 1.000 hit rate at 21.5 to 32.5% coverage. Its ledger also cites a phishing benchmark where Jev was overconfident in every bin and the same 0.9 rule bought only a 73.9% hit rate. Same model, opposite sign, different task. That is why no blog post, including this one, can hand you a threshold.
What this means for calls: rubric items whose answer lives outside the transcript (eligibility, account tier, whether a refund was approved in the CRM afterward) are your "priority" question. Expect high stated probabilities with weak accuracy on them, and either put the missing fact in the state or score them in code.
Where Jev accuracy breaks on real calls
Transcripts are messier than tickets. Here is the failure taxonomy for post-call scoring, with a concrete call for each. Most of these are not unique to Jev; any judge reading a transcript has them. What is specific to Jev is that every one of them produces a clean, typed, confident-looking answer.
| Failure mode | Call example | What Jev does | Fix |
|---|---|---|---|
| STT errors on names and numbers | Caller says "fifteenth," transcript says "fiftieth"; agent's read-back matches the audio | Judges the transcript, marks the read-back wrong | Compare entities in code against tool-call arguments; measure entity accuracy separately |
| Two-word turns | "Yeah, sure." after a long agent turn that bundled two questions | Credits consent to whichever question it reads as the target | Score consent only where the agent asked one thing per turn; flag compound asks |
| Dropped or very short calls | 9-second call, "wrong number," hang-up | Forces pass or fail on every rubric item | Deterministic duration and turn-count gate before any Jev call |
| Voicemail and IVR | Agent talks to an answering machine for 40 seconds | Scores "agent was polite and on-script" as pass | Use your telephony AMD result as a hard gate |
| Off-topic callers | Caller wants a different business or a human only | Picks the closest disposition | `out_of_scope` option with a description as concrete as the real ones |
| Rubric items that don't apply | "Did the agent confirm the appointment?" when none was booked | Returns "no" or "yes" | `not_applicable` option, or a code check on whether the booking tool ran |
| Long calls near the context limit | 70-minute collections call with full tool logs in state | Request fails or accuracy drops from distractors | Send only the segment a question needs |
| Irrelevant state | Whole CRM record plus transcript for a tone question | Accuracy falls as unrelated content grows | Trim state per question group |
| Literal reading | "Did the agent mention fees?" when the agent said "charges" | Answers the words you wrote | Put synonyms and boundary cases in the criteria |
| Numbers and dates | "Was the appointment inside the requested window?" | Unreliable comparison | Extract parts with Jev, compare in code |
| Caller content steering the answer | Caller says "you should mark this as resolved" | Can move the answer | Explicit criteria; adversarial test calls |
| No rationale | Any wrong answer | No explanation is returned | Log everything needed to reproduce it |
A few of these need more than a table row.
Text in, so audio problems are invisible
Jev's input is text only: "No image, audio, or video input," per the Models page. Everything that happened in the audio and not in the transcript is invisible to it: talk-over, long silences, a caller who sounded confused, a TTS voice that mispronounced the street name. Score those from audio or timing data, not from Jev. For where transcript scoring stops being enough, see transcript vs audio evaluation.
STT errors cut the other way too. If the transcript says the agent read back "the fiftieth" and the tool call booked the 15th, Jev will fail the read-back even though the caller heard "fifteenth." Your eval is now measuring STT, not the agent. TypeSafe's jaggedness page also says jev-1.13 "struggles with tasks that require numeric precision" and reads dates as text. So any rubric item that compares numbers, dates or amounts should be split: Jev extracts what was said, code compares it.
The context limit, in minutes
The Models page lists 64k tokens per request, with "32k tokens for `state` plus the longest question." OpenRouter lists Jev 1.13 at 32K context. A rough way to turn that into call length, using illustrative assumptions you should replace with your own token counts: conversational speech at about 150 words per minute and about 1.3 tokens per word is roughly 200 tokens per minute of plain transcript. That suggests 32k covers about 160 minutes of bare text. But voice pipelines rarely send bare text. Speaker labels, timestamps per turn, JSON structure and tool-call payloads can easily double or triple the token count, which pulls the practical ceiling down to roughly 50 to 80 minutes.
The more important limit comes earlier. The jaggedness page says "Accuracy falls as the state grows with content unrelated to the decision" and recommends filtering first and sending "only the fields the question needs." A tone question does not need the payment tool logs. An identity-verification question needs the first two minutes, not the whole call.
Questions in one request can't see each other
TypeSafe's build guide says "Questions are evaluated independently and in parallel. One primitive's result does not become hidden context that changes another primitive's result." jev-exploration tested it: a code word placed in a sibling question's instructions was acted on at 0.02, versus 0.99 when the same code word was in the shared state.
That rules out rubric logic like "If an appointment was booked, did the agent confirm it?" inside one question, with a separate "Was an appointment booked?" question next to it. The second question's answer is not visible to the first. Ask both as independent questions, then combine them in code: `confirmed` only counts when `booked` is true.
No rationale, so log what makes it reproducible
Jev returns no explanation. When an answer is wrong, the only way to debug it is to rerun the exact request. Log the question text, every option description, the state exactly as sent, the full probability vector and the `model` field from the response (for example `jev-1.13.0`). The `jev-latest` alias moves when a new version ships, and the Models page recommends pinning the versioned ID if you have tuned thresholds.
Option order
TypeSafe documents it: "the order of a Choice's options can affect the answer, and jev-1.13 leans toward the option that comes first." A Hacker News commenter made the same point more sharply in the launch thread, noting that reordering choices can shift probabilities and that the `confidence` output is a formula over those probabilities. This is not unique to Jev. Zheng et al. found LLMs prefer particular option positions or IDs in multiple-choice questions. For scoring, the practical risk is putting `pass` first in every rubric item. Test it with a flip run (covered in the How to section below).
What forced answers do to your QA metrics
Here is the part that turns a modeling detail into a business problem. The numbers below are a labeled, illustrative example. Plug in your own call mix.
Assume 10,000 calls a week. 8% (800) are short or dropped: hang-ups in the first seconds, wrong numbers, voicemail that slipped past AMD. The other 9,200 are real conversations.
Example 1: a compliance rate
The rubric item is "Agent read the recording disclosure." Assume the true compliance rate on real calls is 98%. On short calls the disclosure was often never due, because the caller hung up before the agent finished its greeting. With no escape option, assume Jev fails 90% of those short calls (it sees no disclosure, so "no" is the least-wrong answer).
- True compliance: 9,200 × 0.98 = 9,016 compliant of 9,200 = 98.0%
- Short calls forced: 800 × 0.10 = 80 pass, 720 fail
- Reported: (9,016 + 80) / 10,000 = 91.0%
A seven-point compliance drop that does not exist. In a regulated workflow, that number triggers a review.
Now add escape options and assume they catch 90% of short calls (720 abstain, 80 still forced, 8 of those pass):
- Reported: (9,016 + 8) / (9,200 + 80) = 9,024 / 9,280 = 97.2%
Closer, but still 0.8 points low, which is why the escape option alone is not enough. A deterministic gate (duration under 15 seconds or fewer than two caller turns never reaches Jev) catches what the escape option misses.

Example 2: an overall QA pass rate and a fake trend
Rubric item: "Call handled correctly." True pass rate on real calls: 85%. Assume Jev fails 80% of forced short calls.
- Week 1, 8% short calls: (9,200 × 0.85 + 800 × 0.20) / 10,000 = (7,820 + 160) / 10,000 = 79.8% reported vs 85.0% true
- Week 2, a carrier issue pushes short calls to 12%: (8,800 × 0.85 + 1,200 × 0.20) / 10,000 = (7,480 + 240) / 10,000 = 77.2% reported vs 85.0% true
Your agent did not change. Your prompt did not change. The QA score fell 2.6 points because a telephony problem changed the mix of calls Jev was forced to score. Someone will spend a day hunting a prompt regression that does not exist. This is the metric drift you want to catch at the source.
The sign depends on how the rubric item is worded
Notice that both examples moved the metric down. That is because both items are phrased as "did X happen," and on a call with no content, "no" is the least-wrong answer. Flip the wording to an absence ("Agent avoided prohibited language," "Agent made no unapproved promises") and empty calls drift toward "pass," which inflates the metric instead. A rubric with both kinds of items gets pushed in both directions at once, and the averages can hide it. The only clean answer is to remove non-applicable calls from the denominator.
Seven fixes for Jev limitations in call scoring
These come straight from the mechanism above. Each one is cheap.
1. Put an escape option on every Choice and Score. Use two: `not_applicable` (the situation never arose) and `insufficient_evidence` (the transcript is too short, garbled or cut off to tell). Write them as concretely as the real options. "Other" is too vague; "No appointment was discussed, offered or booked on this call" gives the model something to match. For a Score, add the escape as a separate Choice question, since Score levels are ordered and an escape level would sit in the middle of the scale.
2. Run an evidence gate first. Deterministic checks in code come first: call duration, caller turn count, AMD result, whether the relevant tool ran. Then a Noul such as "Does the transcript contain a real conversation in which the caller states a request and the agent responds?" Only calls that pass both get their rubric scores counted.
3. Fit thresholds per question on your own labels. The confidence docs say thresholds "depend on your domain and the performance of the model for your use case." The calibration studies show the sign of the error can differ by question type. One global 0.8 is a guess. Use top probability or a refit calibration map, not the raw `confidence` field.
4. Exclude not_applicable from denominators. Report three numbers per rubric item: the rate on decided calls, coverage (share decided), and the abstain rate. A pass rate without coverage is not interpretable.
5. Track the escape rate as a drift signal. If `insufficient_evidence` jumps from 6% to 11% overnight, something upstream changed: STT, telephony, AMD, a new campaign list. That is often the earliest alarm you will get.
6. Trim state to what the decision needs. Group questions by the evidence they need. Identity verification gets the first turns. Disposition gets the last turns plus tool results. Tone gets caller turns only.
7. Split dependent questions across independent asks. "Was an appointment booked?" and "Did the agent read back the date and time?" are two questions. Code combines them. Never write the condition into one question and hope the model checks the other.
For how these fit a full rubric, see Jev as a judge for voice agents and Jev voice agent evaluation.
Python: escape options, evidence gate, thresholds and honest metrics
This is simplified and illustrative, written against the documented `typesafe-sdk` (`pip install typesafe-sdk`). Names used here (`TypeSafeClient`, `system_one`, `Choice`, `Noul`, `NoulCriteria`, `answers`, `choice`, `probabilities`, `noul`, `model`) match the Python SDK reference and the primitive pages. The client reads `TYPESAFE_API_KEY` from the environment. Thresholds are placeholders until you fit them.
from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient
MODEL = "jev-1.13.0" # pin the version your thresholds were fit on
ESCAPES = {"not_applicable", "insufficient_evidence"}
ESCAPE_CRITERIA = {
"not_applicable": "The situation this question asks about never came up on this call.",
"insufficient_evidence": "The transcript is too short, cut off or garbled to tell either way.",
}
GATE = Noul(
instructions="Does `call.turns` contain a real conversation in which the caller "
"states a request and the agent responds to it?",
criteria=NoulCriteria(
true="The caller says what they want and the agent replies to that request.",
false="Voicemail, silence, wrong number, a hang-up, or only greetings.",
),
)
RUBRIC = {
"appointment_booked": Choice(
instructions="Was an appointment booked on this call?",
criteria={
"yes": "The agent and caller agreed on a specific date and time.",
"no": "An appointment was discussed but none was agreed.",
**ESCAPE_CRITERIA,
},
),
"readback_done": Choice(
instructions="Before the call ended, did the agent read the booked date "
"and time back to the caller?",
criteria={
"yes": "The agent repeated the date and time to the caller.",
"no": "An appointment was booked but the agent did not repeat it.",
"not_applicable": "No appointment was booked, offered or discussed.",
"insufficient_evidence": ESCAPE_CRITERIA["insufficient_evidence"],
},
),
"disclosure_read": Choice(
instructions="Did the agent say that the call is recorded?",
criteria={
"yes": "The agent told the caller the call is recorded.",
"no": "The conversation went past the greeting and no recording notice was given.",
"not_applicable": "The caller hung up before the agent finished its greeting.",
"insufficient_evidence": ESCAPE_CRITERIA["insufficient_evidence"],
},
),
}
# Fit these per question on your labeled calls. Placeholders only.
THRESHOLDS = {"appointment_booked": 0.80, "readback_done": 0.85, "disclosure_read": 0.90}
GATE_MIN = 0.80
def hard_gate(call: dict) -> str | None:
"""Deterministic checks that never need a model."""
if call["amd_result"] == "machine":
return "voicemail"
if call["duration_s"] < 15 or call["caller_turns"] < 2:
return "too_short"
return None
def route(answer, threshold: float) -> tuple[str, str | None]:
"""Turn a Choice answer into decided / abstain / review."""
probs = dict(answer.probabilities)
escape_mass = sum(probs.get(k, 0.0) for k in ESCAPES)
top_label = max(probs, key=probs.get)
if top_label in ESCAPES:
return "abstain", top_label
if escape_mass >= 0.30: # real doubt even if a real label won
return "review", top_label
if probs[top_label] < threshold: # top probability, not the confidence field
return "review", top_label
return "decided", top_label
def score_call(client: TypeSafeClient, call: dict) -> dict:
reason = hard_gate(call)
if reason:
return {"call_id": call["id"], "gate": reason, "items": {}}
state = {"call": {"turns": call["turns"]}} # trimmed: no CRM blob, no tool logs
questions = {"gate": GATE, **RUBRIC}
resp = client.system_one(model=MODEL, state=state, questions=questions)
record = {
"call_id": call["id"],
"model": resp.model, # log the version that answered
"state": state, # log verbatim so any answer can be replayed
"gate_p": resp.answers["gate"].noul,
"items": {},
}
if record["gate_p"] < GATE_MIN:
record["gate"] = "no_conversation"
return record
for qid in RUBRIC:
status, label = route(resp.answers[qid], THRESHOLDS[qid])
record["items"][qid] = {
"status": status,
"label": label,
"probs": dict(resp.answers[qid].probabilities),
}
# Dependent logic lives in code: a read-back only counts if a booking happened.
booked = record["items"]["appointment_booked"]
if booked["status"] == "decided" and booked["label"] == "no":
record["items"]["readback_done"] = {"status": "abstain", "label": "not_applicable", "probs": {}}
return recordThe gate Noul rides in the same request as the rubric. That is deliberate: TypeSafe says questions run in parallel and adding them "barely changes the response time," so one request is cheaper than two round trips. "Gate first" refers to the order your code reads the answers, not the order of requests.
Now the metric code. This is where most dashboards go wrong.
from collections import Counter
def item_metrics(records: list[dict], qid: str, pass_label: str = "yes") -> dict:
total = len(records)
statuses = Counter()
passes = 0
for r in records:
if "gate" in r and r["gate"]:
statuses["gated_out"] += 1
continue
item = r["items"].get(qid)
statuses[item["status"]] += 1
if item["status"] == "decided" and item["label"] == pass_label:
passes += 1
decided = statuses["decided"]
return {
"pass_rate": passes / decided if decided else None, # denominator excludes N/A
"coverage": decided / total if total else None,
"abstain_rate": statuses["abstain"] / total if total else None, # drift signal
"review_rate": statuses["review"] / total if total else None,
"gated_out_rate": statuses["gated_out"] / total if total else None,
}Put `abstain_rate` and `gated_out_rate` on the same chart as the pass rate. If pass rate drops and abstain rate climbs on the same day, look at telephony and STT before you look at the prompt.

Live-call decisions: the same rule, with a different escape
Most of this post is about scoring calls after they end. The same failure shows up during the call when Jev drives routing or tool gating, and it costs more, because a forced answer becomes an action.
Example: after the caller says "yeah no, the other one," your agent asks Jev which of three intents to route to. Two words carry almost no evidence. Without an escape, Jev picks one of the three. With escapes, the options include `ask_clarifying_question` ("The caller's last turn does not say which request they mean") and `route_to_human` ("The caller asked for a person, or the request fits none of the listed intents"). Your code then asks a follow-up instead of firing the wrong flow.
The thresholds should also track stakes, as the confidence docs show with a balance check (low stakes) next to a transfer approval (high stakes, higher bar). For a refund or payment tool, a borderline answer should mean "confirm with the caller," not "run the tool." Latency is not the obstacle: OpenRouter's telemetry for Jev showed a P50 of 0.21 s and a P95 of 0.34 s as of October 1, 2026, and TypeSafe quotes 70 to 500 ms. Budget for the P95 on your own network path, not the median. For the full design, see Jev for routing and escalation and Jev for tool calling.
How to test Jev accuracy on your own calls
The full benchmark protocol, with sample-size math, kappa and the latency harness, is in How to benchmark a decision model on your own call data. Here is the shorter version focused on the abstain problem.
1. Pull a stratified sample of 300 to 500 real calls. Oversample the cases that break forced choice: calls under 20 seconds, voicemails, wrong numbers, transfers, calls over 30 minutes, non-English callers and calls with known STT trouble (names, addresses, account numbers). A random sample will under-represent exactly the calls you need.
2. Label gold answers including "not applicable" and "insufficient evidence." Every rubric item gets one of its real labels or an escape label. Have two people label 20% of the set to measure your own agreement. If humans disagree on an item, Jev cannot be held to a higher bar on it. The golden dataset guide covers labeling rules.
3. Run Jev with escape options and a pinned model version. Store the full probability vector for every answer, the state as sent and the `model` field.
4. Measure escape precision and recall per item. Recall: of calls labeled not applicable or insufficient, what share did Jev route to an escape? Precision: of Jev's escapes, what share were truly not applicable? Low recall means forced answers are still leaking into your metrics. Low precision means real calls are being thrown out.
5. Plot coverage against accuracy at several thresholds. For each item, compute accuracy on answered calls and the share answered at top-probability cutoffs of 0.5, 0.6, 0.7, 0.8 and 0.9. Pick the threshold that meets your error budget for that item. A disclosure item may need 0.9; a tone item may be fine at 0.6.
6. Check calibration on your data. Bin predictions, compare stated probability to observed accuracy, and report ECE next to its noise floor for your sample size. If it is off, fit a temperature or Platt map per question. jev-exploration found about 50 labels was the floor for refitting the intercept of such a map on its data.
7. Run an option-order flip test. Reverse or shuffle the option order on the same calls and count how often the selected label changes. Any item with a meaningful flip rate needs better option descriptions or an ensemble over two orders.
8. Run a distractor test. Score the same calls with the trimmed state and with the full state (CRM record, tool logs). The accuracy gap tells you how much state discipline is worth for your pipeline.
9. Repeat the run three times. Jev is highly consistent, but TypeSafe's own cookbook still saw labels flip on borderline questions. Report the flip rate per item and treat flipping items as review candidates.
10. Re-run when anything upstream changes. New STT model, new TTS voice, new prompt, new Jev version behind `jev-latest`. The labeled set becomes your regression suite.
How independent evaluation fits
Running Jev is easy. Knowing whether its answers are right on your calls takes labeled data, and labeling is where in-house teams stall. Evalgent is an independent evaluator: we label real and synthetic calls against your rubric, including the short, dropped and off-topic calls that break forced choice, and report escape precision and recall, coverage curves and calibration per item.
That is useful at three points: before you trust Jev scores for a compliance report, when you compare Jev with an LLM judge or OpenAI's Decisions API, and as a regression check when your STT, prompt or model version changes. For the broader trade-offs of model judges, see the limits of LLM-as-judge for voice AI. For always-on use, see Jev for call monitoring.
Frequently asked questions
How accurate is Jev for scoring voice agent calls?
No one has published Jev accuracy on call transcripts. LangChain's study found 500 of 500 pass/fail matches on five fixed agent runs. Independent tests range from 89 to 92% on clear ticket questions to 44.7% on a rule hidden from the text. Accuracy depends on the task, so measure it on your own labeled calls.
Can Jev say "I don't know"?
Not by itself. A Choice always returns one of the options you supplied, and a Noul always returns a probability of yes. If no option fits, Jev picks the least-wrong one. You get abstention by adding explicit escape options such as not_applicable and insufficient_evidence, and by routing low top probabilities to review in code.
Is the Jev confidence score a probability of being correct?
No. TypeSafe computes Choice confidence from the shape of the probability distribution: (p_max − 1/n) / (1 − 1/n). It measures how peaked the answer is, not whether it is right. One independent study found it was never better than the top probability as a correctness signal. Threshold on calibrated top probability per question instead.
Why do short or dropped calls hurt my QA pass rate?
Jev must answer every rubric question on every call. On a nine-second call, "did the agent read the disclosure" usually comes back "no" because nothing happened. Those forced fails enter your denominator. In an illustrative 10,000-call week with 8% short calls, a true 98% compliance rate reports as 91%.
What escape options should I add to a Jev Choice question?
Use not_applicable for situations that never arose, such as no appointment booked, and insufficient_evidence for transcripts that are too short, cut off or garbled. Describe each as concretely as your real options. For live-call routing, use route_to_human and ask_clarifying_question so a weak answer becomes a safe action.
Does Jev work on long calls?
Within limits. TypeSafe lists 64k tokens per request and 32k for the state plus the longest question; OpenRouter lists 32K. With speaker labels, timestamps and tool logs, that can be well under two hours of talk. Accuracy also falls as irrelevant state grows, so send each question only the segment it needs.
Do Jev questions in the same request see each other's answers?
No. TypeSafe states that questions are evaluated independently and in parallel, and one answer does not become context for another. An independent probe confirmed it. Ask "was an appointment booked" and "did the agent read it back" separately, then combine them in code.
Does Jev give the same answer every time?
Not always. Jev is highly consistent, much more so than LLM judges in LangChain's study, but answers on borderline items can still change. TypeSafe's own cookbook saw labels flip on 2 of 8 borderline questions across 15 runs. Repeat your evaluation runs, track flip rates per rubric item and send flipping items to review.
The bottom line
Jev accuracy is good enough to score most voice agent decisions, and its consistency and cost make full-coverage scoring practical, but it will always answer, even when a call gives it nothing to answer with. Build the abstention yourself with escape options, evidence gates, per-question thresholds and honest denominators, then prove it on your own labeled calls before your dashboard drives decisions.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more