Evalgent
Back to Blog
Voice AI Evaluation

Jev Accuracy in Voice Agent Evals: Where It Fails on Calls

Deepesh Jayal
23 min read
Jev Accuracy in Voice Agent Evals: Where It Fails on Calls
On this page

You wire Jev into your post-call pipeline. Every call gets a QA rubric: did the agent verify identity, read the recording disclosure, confirm the appointment, resolve the issue. Jev answers each question in a fraction of a second for a fraction of a cent. The dashboard fills up.

Then someone listens to a call that Jev failed on "confirmed the appointment." It lasted nine seconds. The caller said "wrong number" and hung up. There was no appointment to confirm. Jev did not make a typo or hallucinate a new label. It did what it is built to do: pick one of the options you gave it.

This guide answers one question for teams that run their own voice agent: can you trust Jev to score your calls? The answer is yes for a large share of decisions, provided you design around one structural limitation. We cover what the public accuracy evidence actually shows, where Jev breaks on real call transcripts, what forced answers do to your pass rates, the fixes, working Python against the documented SDK, and how to test all of it on your own calls.

500/500
Jev agreement with human labels on pass/fail in LangChain's 5-case, 100-repeat study
0.52 at 0.04
Winning probability and confidence when Jev was forced to pick between two wrong options (jev-exploration, single probe)
0.74 vs 44.7%
Average stated probability vs accuracy on a rule Jev could not know (jev-ood-calibration)

What the evidence says about Jev accuracy

Jev is TypeSafe's first System One model. It takes a state (your transcript plus metadata) and typed questions, and returns answers with probabilities instead of generated text. There are three question types: Choice (pick one option), Score (a position on ordered levels) and Noul (the probability a yes/no statement is true). If that is new, start with Jev for voice agents.

Four kinds of evidence exist so far. They measure different things, so read them side by side.

SourceWho ran itTask and sizeHeadline resultWhat it cannot tell you
LangChain, Jev-as-a-JudgeLangChain team5 fixed weather-agent runs, 100 repeats each (500 decisions)Pass/fail matched the human oracle on 500 of 500; GPT-5.6 Terra 99.8%, GPT-5.6 Luna 96.4%, Claude Sonnet 4.6 80.0%. Quality-score variance 92 to 913 times lower than the LLM judges. $0.34 total vs $28.17 for Claude; 0.44 s and $0.00035 per callOnly five unique cases, so agreement is about repeatability on a narrow task. LLM judges ran at provider defaults; the Jev version was not recorded
TypeSafe launch postTypeSafeFour internal "workflow evals"193.6x faster and 444.6x cheaper than frontier LLMs. TypeSafe itself writes: "we expect that these are on the higher end of real world gains"Reference answers are the average of two frontier LLMs, not human ground truth (TypeSafe says so)
jev-explorationOne independent tester12 support tickets; 4 edge-case probesDepartment routing 10 of 12 (11 of 12 after a label revision), urgency 12 of 12n=12 and n=4. The author says the 95% intervals overlap a trivial baseline (majority class got 5 of 12), so this is not accuracy evidence
jev-ood-calibrationOne independent tester900 synthetic support tickets Jev cannot have seen, 3 question typesQueue 89.0%, "is the customer angry" 91.7%, priority (a hidden policy rule) 44.7%Synthetic text, one task family. Tells you about calibration out of distribution, not about call transcripts

Credit where it is due. Every study that measured it found Jev highly consistent. LangChain's variance gap is large. TypeSafe's own self-consistency cookbook ran an 8-question moderation rubric 15 times: raw label agreement was 90.8%, and with a rule that sends any answer under 0.60 top probability to a human, agreement rose to 99.2% with 74.2% of answers handled automatically. The cost is low enough that you can score every call instead of a sample.

Two caveats run through all of it. First, consistency is not correctness. A judge that is reliably wrong produces bad data at scale, and LangChain says exactly that in its write-up. Second, none of the studies above used call transcripts. Accuracy swings by task in the independent ledger that jev-exploration maintains, from 98.3% on spam to 62.6% on phishing. Your calls are a task nobody has published.

The core failure: Jev cannot say "I don't know"

This is the mechanism to understand before you trust any number Jev gives you.

A Choice answer is a probability distribution over the options you supplied. The probabilities sum to 1. The `choice` field is "the option with the highest probability," per the Choice docs. The model cannot return an option you did not offer. So if none of your options fit the call, Jev still returns one of them. It picks the least-wrong one.

That is the price of type safety. TypeSafe's launch post is candid about it: its 0% hallucination figure carries the footnote "Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots." A valid label is guaranteed. A true label is not. TypeSafe's System One page says the same about probabilities: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct."

What the edge-case probes showed

The jev-exploration author ran four probes built to have no good answer. These are single probes from one tester, and the raw responses were not kept, so treat them as illustrations of behavior, not rates:

ProbeWhat was askedJev's output
Forced Choice, no valid optionPick between two options that were both wrongWinner at 0.52, confidence 0.04
Unknowable Noul"Will the user renew?" with no evidence either way0.49
Contradictory evidenceState supports both yes and no0.40
Question unrelated to stateAsked about something not in the stateAnswered from world knowledge at 0.04

The first row is the one that bites call scoring. TypeSafe computes Choice confidence as (p_max − 1/n) / (1 − 1/n), per the confidence docs. With two options, 0.52 works out to (0.52 − 0.5) / 0.5 = 0.04. The distribution did flag the problem. The answer did not. If your pipeline stores only `choice`, your database says "pass" or "fail" with nothing to show it was a coin flip.

The fourth row is worse for Noul questions. A Noul has no third outcome. "Did the agent verify the caller's identity?" on a six-second call returns a value near 0, which looks exactly like a confident "no." It is literally true that no verification happened. It is wrong for a compliance metric, because verification was never due.

Two bar charts compare a forced two-option Jev Choice that logs pass at 0.52 probability with 0.04 confidence against an illustrative four-option Choice where insufficient_evidence takes most probability

Why this is an old problem with a known fix

Machine reading had the same failure. The SQuAD 2.0 paper, "Know What You Don't Know", added over 50,000 unanswerable questions to a benchmark where every question used to have an answer. A strong system that scored 86% F1 on the original set scored 66% F1 once it also had to decide when to abstain. Evaluation sets without "no answer" cases overstate accuracy. Your QA rubric is the same: if your test calls all have an appointment, you never see how the "appointment confirmed" item behaves on calls without one.

The fix is equally old. Selective classification, or the reject option, lets a model decline when it is not sure, trading coverage for lower error. Geifman and El-Yaniv showed you can pick a target error rate and reject inputs until you hit it; on ImageNet they guaranteed 2% top-5 error at almost 60% coverage. Jev gives you the raw material (a full distribution) but not the reject decision. You build it two ways: an explicit escape option the model can choose, and a threshold in your code. TypeSafe recommends the first one directly: "Add an `other` or `none of the above` option when the list might not cover every input, so the model can say none of the others fit."

Jev confidence score calibration when calls look unfamiliar

The second half of the abstain problem is whether the probabilities warn you. The best public test of that is jev-ood-calibration, which generated 900 support tickets from a script so Jev could not have seen them. One of three questions, ticket priority, depended on a policy rule that is not in the text (customer tier bumps priority). No model can recover that from the ticket. It is the closest public analog to a rubric item that depends on something outside your transcript.

What it found, as published on September 19, 2026:

MeasureResultWhat it means
Accuracy on the unknowable priority rule44.7% (chance is 25%)Common sense gets partway; the rule itself is invisible
Average stated probability on that question0.74Jev did not signal that it could not know
ECE across all 900 items0.107, against a noise floor of 0.024 (4.4x)Real miscalibration, not sampling noise
Refit temperature, all 9002.74Probabilities were overconfident overall
Refit temperature by typeChoice 3.29, Score 3.40, boolean 0.66Choice and Score overconfident; booleans underconfident
Exact zeros (OpenBookQA run)1,051 of 2,000 option probabilities were exactly 0.00One item put 0.00 on the correct answer; no temperature fix can repair a hard zero

The temperature refit is the method from Guo et al., "On Calibration of Modern Neural Networks": find the single scalar that makes probabilities honest after the fact. T above 1 means the model is overconfident, below 1 means underconfident. The author's practical advice follows from the sign flip: calibrate per question, not per model, and do not threshold on the `confidence` field, which was never better than the top probability in that study and sometimes much worse.

jev-exploration adds two more findings. Every probability Jev returned in its 281 committed responses sat on a 0.01 grid. And on its own 800-item set, Noul probabilities were compressed toward the middle, so a 0.9 threshold gave a 1.000 hit rate at 21.5 to 32.5% coverage. Its ledger also cites a phishing benchmark where Jev was overconfident in every bin and the same 0.9 rule bought only a 73.9% hit rate. Same model, opposite sign, different task. That is why no blog post, including this one, can hand you a threshold.

What this means for calls: rubric items whose answer lives outside the transcript (eligibility, account tier, whether a refund was approved in the CRM afterward) are your "priority" question. Expect high stated probabilities with weak accuracy on them, and either put the missing fact in the state or score them in code.

Where Jev accuracy breaks on real calls

Transcripts are messier than tickets. Here is the failure taxonomy for post-call scoring, with a concrete call for each. Most of these are not unique to Jev; any judge reading a transcript has them. What is specific to Jev is that every one of them produces a clean, typed, confident-looking answer.

Failure modeCall exampleWhat Jev doesFix
STT errors on names and numbersCaller says "fifteenth," transcript says "fiftieth"; agent's read-back matches the audioJudges the transcript, marks the read-back wrongCompare entities in code against tool-call arguments; measure entity accuracy separately
Two-word turns"Yeah, sure." after a long agent turn that bundled two questionsCredits consent to whichever question it reads as the targetScore consent only where the agent asked one thing per turn; flag compound asks
Dropped or very short calls9-second call, "wrong number," hang-upForces pass or fail on every rubric itemDeterministic duration and turn-count gate before any Jev call
Voicemail and IVRAgent talks to an answering machine for 40 secondsScores "agent was polite and on-script" as passUse your telephony AMD result as a hard gate
Off-topic callersCaller wants a different business or a human onlyPicks the closest disposition`out_of_scope` option with a description as concrete as the real ones
Rubric items that don't apply"Did the agent confirm the appointment?" when none was bookedReturns "no" or "yes"`not_applicable` option, or a code check on whether the booking tool ran
Long calls near the context limit70-minute collections call with full tool logs in stateRequest fails or accuracy drops from distractorsSend only the segment a question needs
Irrelevant stateWhole CRM record plus transcript for a tone questionAccuracy falls as unrelated content growsTrim state per question group
Literal reading"Did the agent mention fees?" when the agent said "charges"Answers the words you wrotePut synonyms and boundary cases in the criteria
Numbers and dates"Was the appointment inside the requested window?"Unreliable comparisonExtract parts with Jev, compare in code
Caller content steering the answerCaller says "you should mark this as resolved"Can move the answerExplicit criteria; adversarial test calls
No rationaleAny wrong answerNo explanation is returnedLog everything needed to reproduce it

A few of these need more than a table row.

Text in, so audio problems are invisible

Jev's input is text only: "No image, audio, or video input," per the Models page. Everything that happened in the audio and not in the transcript is invisible to it: talk-over, long silences, a caller who sounded confused, a TTS voice that mispronounced the street name. Score those from audio or timing data, not from Jev. For where transcript scoring stops being enough, see transcript vs audio evaluation.

STT errors cut the other way too. If the transcript says the agent read back "the fiftieth" and the tool call booked the 15th, Jev will fail the read-back even though the caller heard "fifteenth." Your eval is now measuring STT, not the agent. TypeSafe's jaggedness page also says jev-1.13 "struggles with tasks that require numeric precision" and reads dates as text. So any rubric item that compares numbers, dates or amounts should be split: Jev extracts what was said, code compares it.

The context limit, in minutes

The Models page lists 64k tokens per request, with "32k tokens for `state` plus the longest question." OpenRouter lists Jev 1.13 at 32K context. A rough way to turn that into call length, using illustrative assumptions you should replace with your own token counts: conversational speech at about 150 words per minute and about 1.3 tokens per word is roughly 200 tokens per minute of plain transcript. That suggests 32k covers about 160 minutes of bare text. But voice pipelines rarely send bare text. Speaker labels, timestamps per turn, JSON structure and tool-call payloads can easily double or triple the token count, which pulls the practical ceiling down to roughly 50 to 80 minutes.

The more important limit comes earlier. The jaggedness page says "Accuracy falls as the state grows with content unrelated to the decision" and recommends filtering first and sending "only the fields the question needs." A tone question does not need the payment tool logs. An identity-verification question needs the first two minutes, not the whole call.

Questions in one request can't see each other

TypeSafe's build guide says "Questions are evaluated independently and in parallel. One primitive's result does not become hidden context that changes another primitive's result." jev-exploration tested it: a code word placed in a sibling question's instructions was acted on at 0.02, versus 0.99 when the same code word was in the shared state.

That rules out rubric logic like "If an appointment was booked, did the agent confirm it?" inside one question, with a separate "Was an appointment booked?" question next to it. The second question's answer is not visible to the first. Ask both as independent questions, then combine them in code: `confirmed` only counts when `booked` is true.

No rationale, so log what makes it reproducible

Jev returns no explanation. When an answer is wrong, the only way to debug it is to rerun the exact request. Log the question text, every option description, the state exactly as sent, the full probability vector and the `model` field from the response (for example `jev-1.13.0`). The `jev-latest` alias moves when a new version ships, and the Models page recommends pinning the versioned ID if you have tuned thresholds.

Option order

TypeSafe documents it: "the order of a Choice's options can affect the answer, and jev-1.13 leans toward the option that comes first." A Hacker News commenter made the same point more sharply in the launch thread, noting that reordering choices can shift probabilities and that the `confidence` output is a formula over those probabilities. This is not unique to Jev. Zheng et al. found LLMs prefer particular option positions or IDs in multiple-choice questions. For scoring, the practical risk is putting `pass` first in every rubric item. Test it with a flip run (covered in the How to section below).

What forced answers do to your QA metrics

Here is the part that turns a modeling detail into a business problem. The numbers below are a labeled, illustrative example. Plug in your own call mix.

Assume 10,000 calls a week. 8% (800) are short or dropped: hang-ups in the first seconds, wrong numbers, voicemail that slipped past AMD. The other 9,200 are real conversations.

Example 1: a compliance rate

The rubric item is "Agent read the recording disclosure." Assume the true compliance rate on real calls is 98%. On short calls the disclosure was often never due, because the caller hung up before the agent finished its greeting. With no escape option, assume Jev fails 90% of those short calls (it sees no disclosure, so "no" is the least-wrong answer).

  • True compliance: 9,200 × 0.98 = 9,016 compliant of 9,200 = 98.0%
  • Short calls forced: 800 × 0.10 = 80 pass, 720 fail
  • Reported: (9,016 + 80) / 10,000 = 91.0%

A seven-point compliance drop that does not exist. In a regulated workflow, that number triggers a review.

Now add escape options and assume they catch 90% of short calls (720 abstain, 80 still forced, 8 of those pass):

  • Reported: (9,016 + 8) / (9,200 + 80) = 9,024 / 9,280 = 97.2%

Closer, but still 0.8 points low, which is why the escape option alone is not enough. A deterministic gate (duration under 15 seconds or fewer than two caller turns never reaches Jev) catches what the escape option misses.

Waterfall chart shows an illustrative compliance rate falling from a true 98.0 percent to a reported 91.0 percent when short calls are forced into fail, recovering to 97.2 percent with escape options

Example 2: an overall QA pass rate and a fake trend

Rubric item: "Call handled correctly." True pass rate on real calls: 85%. Assume Jev fails 80% of forced short calls.

  • Week 1, 8% short calls: (9,200 × 0.85 + 800 × 0.20) / 10,000 = (7,820 + 160) / 10,000 = 79.8% reported vs 85.0% true
  • Week 2, a carrier issue pushes short calls to 12%: (8,800 × 0.85 + 1,200 × 0.20) / 10,000 = (7,480 + 240) / 10,000 = 77.2% reported vs 85.0% true

Your agent did not change. Your prompt did not change. The QA score fell 2.6 points because a telephony problem changed the mix of calls Jev was forced to score. Someone will spend a day hunting a prompt regression that does not exist. This is the metric drift you want to catch at the source.

The sign depends on how the rubric item is worded

Notice that both examples moved the metric down. That is because both items are phrased as "did X happen," and on a call with no content, "no" is the least-wrong answer. Flip the wording to an absence ("Agent avoided prohibited language," "Agent made no unapproved promises") and empty calls drift toward "pass," which inflates the metric instead. A rubric with both kinds of items gets pushed in both directions at once, and the averages can hide it. The only clean answer is to remove non-applicable calls from the denominator.

Find out where your call scoring is wrong before your dashboard does
An independent audit labels a sample of your real calls, including the short, dropped and off-topic ones, and measures how Jev or any judge behaves on them.
Book a demo

Seven fixes for Jev limitations in call scoring

These come straight from the mechanism above. Each one is cheap.

1. Put an escape option on every Choice and Score. Use two: `not_applicable` (the situation never arose) and `insufficient_evidence` (the transcript is too short, garbled or cut off to tell). Write them as concretely as the real options. "Other" is too vague; "No appointment was discussed, offered or booked on this call" gives the model something to match. For a Score, add the escape as a separate Choice question, since Score levels are ordered and an escape level would sit in the middle of the scale.

2. Run an evidence gate first. Deterministic checks in code come first: call duration, caller turn count, AMD result, whether the relevant tool ran. Then a Noul such as "Does the transcript contain a real conversation in which the caller states a request and the agent responds?" Only calls that pass both get their rubric scores counted.

3. Fit thresholds per question on your own labels. The confidence docs say thresholds "depend on your domain and the performance of the model for your use case." The calibration studies show the sign of the error can differ by question type. One global 0.8 is a guess. Use top probability or a refit calibration map, not the raw `confidence` field.

4. Exclude not_applicable from denominators. Report three numbers per rubric item: the rate on decided calls, coverage (share decided), and the abstain rate. A pass rate without coverage is not interpretable.

5. Track the escape rate as a drift signal. If `insufficient_evidence` jumps from 6% to 11% overnight, something upstream changed: STT, telephony, AMD, a new campaign list. That is often the earliest alarm you will get.

6. Trim state to what the decision needs. Group questions by the evidence they need. Identity verification gets the first turns. Disposition gets the last turns plus tool results. Tone gets caller turns only.

7. Split dependent questions across independent asks. "Was an appointment booked?" and "Did the agent read back the date and time?" are two questions. Code combines them. Never write the condition into one question and hope the model checks the other.

For how these fit a full rubric, see Jev as a judge for voice agents and Jev voice agent evaluation.

Python: escape options, evidence gate, thresholds and honest metrics

This is simplified and illustrative, written against the documented `typesafe-sdk` (`pip install typesafe-sdk`). Names used here (`TypeSafeClient`, `system_one`, `Choice`, `Noul`, `NoulCriteria`, `answers`, `choice`, `probabilities`, `noul`, `model`) match the Python SDK reference and the primitive pages. The client reads `TYPESAFE_API_KEY` from the environment. Thresholds are placeholders until you fit them.

from typesafe_sdk import Choice, Noul, NoulCriteria, TypeSafeClient

MODEL = "jev-1.13.0"  # pin the version your thresholds were fit on
ESCAPES = {"not_applicable", "insufficient_evidence"}

ESCAPE_CRITERIA = {
    "not_applicable": "The situation this question asks about never came up on this call.",
    "insufficient_evidence": "The transcript is too short, cut off or garbled to tell either way.",
}

GATE = Noul(
    instructions="Does `call.turns` contain a real conversation in which the caller "
                 "states a request and the agent responds to it?",
    criteria=NoulCriteria(
        true="The caller says what they want and the agent replies to that request.",
        false="Voicemail, silence, wrong number, a hang-up, or only greetings.",
    ),
)

RUBRIC = {
    "appointment_booked": Choice(
        instructions="Was an appointment booked on this call?",
        criteria={
            "yes": "The agent and caller agreed on a specific date and time.",
            "no": "An appointment was discussed but none was agreed.",
            **ESCAPE_CRITERIA,
        },
    ),
    "readback_done": Choice(
        instructions="Before the call ended, did the agent read the booked date "
                     "and time back to the caller?",
        criteria={
            "yes": "The agent repeated the date and time to the caller.",
            "no": "An appointment was booked but the agent did not repeat it.",
            "not_applicable": "No appointment was booked, offered or discussed.",
            "insufficient_evidence": ESCAPE_CRITERIA["insufficient_evidence"],
        },
    ),
    "disclosure_read": Choice(
        instructions="Did the agent say that the call is recorded?",
        criteria={
            "yes": "The agent told the caller the call is recorded.",
            "no": "The conversation went past the greeting and no recording notice was given.",
            "not_applicable": "The caller hung up before the agent finished its greeting.",
            "insufficient_evidence": ESCAPE_CRITERIA["insufficient_evidence"],
        },
    ),
}

# Fit these per question on your labeled calls. Placeholders only.
THRESHOLDS = {"appointment_booked": 0.80, "readback_done": 0.85, "disclosure_read": 0.90}
GATE_MIN = 0.80

def hard_gate(call: dict) -> str | None:
    """Deterministic checks that never need a model."""
    if call["amd_result"] == "machine":
        return "voicemail"
    if call["duration_s"] < 15 or call["caller_turns"] < 2:
        return "too_short"
    return None

def route(answer, threshold: float) -> tuple[str, str | None]:
    """Turn a Choice answer into decided / abstain / review."""
    probs = dict(answer.probabilities)
    escape_mass = sum(probs.get(k, 0.0) for k in ESCAPES)
    top_label = max(probs, key=probs.get)
    if top_label in ESCAPES:
        return "abstain", top_label
    if escape_mass >= 0.30:  # real doubt even if a real label won
        return "review", top_label
    if probs[top_label] < threshold:  # top probability, not the confidence field
        return "review", top_label
    return "decided", top_label

def score_call(client: TypeSafeClient, call: dict) -> dict:
    reason = hard_gate(call)
    if reason:
        return {"call_id": call["id"], "gate": reason, "items": {}}

    state = {"call": {"turns": call["turns"]}}  # trimmed: no CRM blob, no tool logs
    questions = {"gate": GATE, **RUBRIC}
    resp = client.system_one(model=MODEL, state=state, questions=questions)

    record = {
        "call_id": call["id"],
        "model": resp.model,  # log the version that answered
        "state": state,       # log verbatim so any answer can be replayed
        "gate_p": resp.answers["gate"].noul,
        "items": {},
    }
    if record["gate_p"] < GATE_MIN:
        record["gate"] = "no_conversation"
        return record

    for qid in RUBRIC:
        status, label = route(resp.answers[qid], THRESHOLDS[qid])
        record["items"][qid] = {
            "status": status,
            "label": label,
            "probs": dict(resp.answers[qid].probabilities),
        }

    # Dependent logic lives in code: a read-back only counts if a booking happened.
    booked = record["items"]["appointment_booked"]
    if booked["status"] == "decided" and booked["label"] == "no":
        record["items"]["readback_done"] = {"status": "abstain", "label": "not_applicable", "probs": {}}
    return record

The gate Noul rides in the same request as the rubric. That is deliberate: TypeSafe says questions run in parallel and adding them "barely changes the response time," so one request is cheaper than two round trips. "Gate first" refers to the order your code reads the answers, not the order of requests.

Now the metric code. This is where most dashboards go wrong.

from collections import Counter

def item_metrics(records: list[dict], qid: str, pass_label: str = "yes") -> dict:
    total = len(records)
    statuses = Counter()
    passes = 0
    for r in records:
        if "gate" in r and r["gate"]:
            statuses["gated_out"] += 1
            continue
        item = r["items"].get(qid)
        statuses[item["status"]] += 1
        if item["status"] == "decided" and item["label"] == pass_label:
            passes += 1
    decided = statuses["decided"]
    return {
        "pass_rate": passes / decided if decided else None,  # denominator excludes N/A
        "coverage": decided / total if total else None,
        "abstain_rate": statuses["abstain"] / total if total else None,  # drift signal
        "review_rate": statuses["review"] / total if total else None,
        "gated_out_rate": statuses["gated_out"] / total if total else None,
    }

Put `abstain_rate` and `gated_out_rate` on the same chart as the pass rate. If pass rate drops and abstain rate climbs on the same day, look at telephony and STT before you look at the prompt.

Pipeline diagram of a post-call scoring flow with code gates, trimmed state, one Jev request containing a Noul evidence gate and Choice questions with escape options, then per-question threshold routing into metrics, review or exclusion

Live-call decisions: the same rule, with a different escape

Most of this post is about scoring calls after they end. The same failure shows up during the call when Jev drives routing or tool gating, and it costs more, because a forced answer becomes an action.

Example: after the caller says "yeah no, the other one," your agent asks Jev which of three intents to route to. Two words carry almost no evidence. Without an escape, Jev picks one of the three. With escapes, the options include `ask_clarifying_question` ("The caller's last turn does not say which request they mean") and `route_to_human` ("The caller asked for a person, or the request fits none of the listed intents"). Your code then asks a follow-up instead of firing the wrong flow.

The thresholds should also track stakes, as the confidence docs show with a balance check (low stakes) next to a transfer approval (high stakes, higher bar). For a refund or payment tool, a borderline answer should mean "confirm with the caller," not "run the tool." Latency is not the obstacle: OpenRouter's telemetry for Jev showed a P50 of 0.21 s and a P95 of 0.34 s as of October 1, 2026, and TypeSafe quotes 70 to 500 ms. Budget for the P95 on your own network path, not the median. For the full design, see Jev for routing and escalation and Jev for tool calling.

How to test Jev accuracy on your own calls

The full benchmark protocol, with sample-size math, kappa and the latency harness, is in How to benchmark a decision model on your own call data. Here is the shorter version focused on the abstain problem.

1. Pull a stratified sample of 300 to 500 real calls. Oversample the cases that break forced choice: calls under 20 seconds, voicemails, wrong numbers, transfers, calls over 30 minutes, non-English callers and calls with known STT trouble (names, addresses, account numbers). A random sample will under-represent exactly the calls you need.

2. Label gold answers including "not applicable" and "insufficient evidence." Every rubric item gets one of its real labels or an escape label. Have two people label 20% of the set to measure your own agreement. If humans disagree on an item, Jev cannot be held to a higher bar on it. The golden dataset guide covers labeling rules.

3. Run Jev with escape options and a pinned model version. Store the full probability vector for every answer, the state as sent and the `model` field.

4. Measure escape precision and recall per item. Recall: of calls labeled not applicable or insufficient, what share did Jev route to an escape? Precision: of Jev's escapes, what share were truly not applicable? Low recall means forced answers are still leaking into your metrics. Low precision means real calls are being thrown out.

5. Plot coverage against accuracy at several thresholds. For each item, compute accuracy on answered calls and the share answered at top-probability cutoffs of 0.5, 0.6, 0.7, 0.8 and 0.9. Pick the threshold that meets your error budget for that item. A disclosure item may need 0.9; a tone item may be fine at 0.6.

6. Check calibration on your data. Bin predictions, compare stated probability to observed accuracy, and report ECE next to its noise floor for your sample size. If it is off, fit a temperature or Platt map per question. jev-exploration found about 50 labels was the floor for refitting the intercept of such a map on its data.

7. Run an option-order flip test. Reverse or shuffle the option order on the same calls and count how often the selected label changes. Any item with a meaningful flip rate needs better option descriptions or an ensemble over two orders.

8. Run a distractor test. Score the same calls with the trimmed state and with the full state (CRM record, tool logs). The accuracy gap tells you how much state discipline is worth for your pipeline.

9. Repeat the run three times. Jev is highly consistent, but TypeSafe's own cookbook still saw labels flip on borderline questions. Report the flip rate per item and treat flipping items as review candidates.

10. Re-run when anything upstream changes. New STT model, new TTS voice, new prompt, new Jev version behind `jev-latest`. The labeled set becomes your regression suite.

How independent evaluation fits

Running Jev is easy. Knowing whether its answers are right on your calls takes labeled data, and labeling is where in-house teams stall. Evalgent is an independent evaluator: we label real and synthetic calls against your rubric, including the short, dropped and off-topic calls that break forced choice, and report escape precision and recall, coverage curves and calibration per item.

That is useful at three points: before you trust Jev scores for a compliance report, when you compare Jev with an LLM judge or OpenAI's Decisions API, and as a regression check when your STT, prompt or model version changes. For the broader trade-offs of model judges, see the limits of LLM-as-judge for voice AI. For always-on use, see Jev for call monitoring.

Frequently asked questions

How accurate is Jev for scoring voice agent calls?

No one has published Jev accuracy on call transcripts. LangChain's study found 500 of 500 pass/fail matches on five fixed agent runs. Independent tests range from 89 to 92% on clear ticket questions to 44.7% on a rule hidden from the text. Accuracy depends on the task, so measure it on your own labeled calls.

Can Jev say "I don't know"?

Not by itself. A Choice always returns one of the options you supplied, and a Noul always returns a probability of yes. If no option fits, Jev picks the least-wrong one. You get abstention by adding explicit escape options such as not_applicable and insufficient_evidence, and by routing low top probabilities to review in code.

Is the Jev confidence score a probability of being correct?

No. TypeSafe computes Choice confidence from the shape of the probability distribution: (p_max − 1/n) / (1 − 1/n). It measures how peaked the answer is, not whether it is right. One independent study found it was never better than the top probability as a correctness signal. Threshold on calibrated top probability per question instead.

Why do short or dropped calls hurt my QA pass rate?

Jev must answer every rubric question on every call. On a nine-second call, "did the agent read the disclosure" usually comes back "no" because nothing happened. Those forced fails enter your denominator. In an illustrative 10,000-call week with 8% short calls, a true 98% compliance rate reports as 91%.

What escape options should I add to a Jev Choice question?

Use not_applicable for situations that never arose, such as no appointment booked, and insufficient_evidence for transcripts that are too short, cut off or garbled. Describe each as concretely as your real options. For live-call routing, use route_to_human and ask_clarifying_question so a weak answer becomes a safe action.

Does Jev work on long calls?

Within limits. TypeSafe lists 64k tokens per request and 32k for the state plus the longest question; OpenRouter lists 32K. With speaker labels, timestamps and tool logs, that can be well under two hours of talk. Accuracy also falls as irrelevant state grows, so send each question only the segment it needs.

Do Jev questions in the same request see each other's answers?

No. TypeSafe states that questions are evaluated independently and in parallel, and one answer does not become context for another. An independent probe confirmed it. Ask "was an appointment booked" and "did the agent read it back" separately, then combine them in code.

Does Jev give the same answer every time?

Not always. Jev is highly consistent, much more so than LLM judges in LangChain's study, but answers on borderline items can still change. TypeSafe's own cookbook saw labels flip on 2 of 8 borderline questions across 15 runs. Repeat your evaluation runs, track flip rates per rubric item and send flipping items to review.

The bottom line

Jev accuracy is good enough to score most voice agent decisions, and its consistency and cost make full-coverage scoring practical, but it will always answer, even when a call gives it nothing to answer with. Build the abstention yourself with escape options, evidence gates, per-question thresholds and honest denominators, then prove it on your own labeled calls before your dashboard drives decisions.

Related Articles