Evalgent
Back to Blog
Voice AI Evaluation

OpenAI Decisions API vs Jev for Voice Agents: What's Known, What Isn't, and How to Choose

Deepesh Jayal
23 min read
OpenAI Decisions API vs Jev for Voice Agents: What's Known, What Isn't, and How to Choose
On this page

Last verified October 4, 2026. OpenAI says broad release is coming "in the coming days." This post will be updated when OpenAI publishes the Decisions API docs and pricing.

Every voice agent makes small, bounded decisions on almost every turn. Which flow does this caller need? Is it safe to run the refund tool? Should a human take over now? What disposition code goes on this call? A general LLM can answer all of these, but it answers slowly and in prose you then have to parse.

Two products now target exactly that slice. TypeSafe launched Jev, a purpose-built decision model, on September 15, 2026. OpenAI announced the Decisions API, powered by a constrained version of GPT-6 Luna, at DevDay on September 29, 2026.

If you run a voice agent in-house on LiveKit, Pipecat or Dograh, you now have to decide whether to wire one in, wait for the other, or build so you can swap. This guide is written from the position of an independent evaluator. Evalgent does not resell either product. The goal is to separate what is confirmed from what is marketing, run the numbers inside a real voice-turn budget, and give you a way to choose this week without regretting it next month.

What is confirmed, what is a vendor claim, and what is not published

Start with the evidence. The table below marks every claim you are likely to see in a comparison post. "Confirmed" means a primary source states it as fact. "Vendor claim" means the vendor said it but conditions are not public or no third party has reproduced it. "Not published" means nobody outside the preview can check it yet.

ClaimStatusSourceDate checked
Decisions API exists; answers user-defined questions with finite predefined answersConfirmedOpenAI DevDay 2026 recapOct 4, 2026
Decisions API is powered by GPT-6 LunaConfirmedOpenAI DevDay recapOct 4, 2026
Decisions API accepts text or image contextConfirmedOpenAI DevDay recapOct 4, 2026
Limited preview; broad release "in the coming days"ConfirmedOpenAI DevDay recap; @OpenAIDevsOct 4, 2026
150 ms vs 1.6 s for the GPT-6 Luna API ("10x faster")Vendor claimDevDay slide reproduced by The DecoderOct 4, 2026
Endpoint path, request and response schema, SDK methodNot publishedNo docs page; guide URL returns no contentOct 4, 2026
Per-option probabilities or confidence returnedNot publishedNoted as unspecified by Vercel and FirecrawlOct 4, 2026
Multiple questions per call, context limitNot publishedFirecrawl docs-index check, Sep 30Oct 4, 2026
Decisions API pricingNot publishedNo row on OpenAI API pricing or the docs pricing tableOct 4, 2026
Standard API key gets 403 "not enabled for this user" on POST /v1/decisionsThird-party reporteesel (not our test)Oct 4, 2026
GPT-6 Luna list price $0.10 / 1M input, $0.50 / 1M outputConfirmed (Luna, not Decisions)GPT-6 Luna model pageOct 4, 2026
Jev endpoint POST api.typesafe.ai/v1/systemone, model `jev-latest`ConfirmedTypeSafe quickstartOct 4, 2026
Jev returns choice, confidence and per-option probabilitiesConfirmedTypeSafe quickstart response bodyOct 4, 2026
Jev price $0.042 / 1M input tokens, output freeConfirmedOpenRouter Jev 1.13Oct 4, 2026
Jev input is text onlyConfirmedTypeSafe docsOct 4, 2026
Jev context 32K on OpenRouter, 64K per TypeSafe docsConfirmed (two figures)OpenRouter; TypeSafe docsOct 4, 2026
Jev latency 70 to 500 msVendor claimTypeSafeOct 4, 2026
Jev p50 0.21 s, p95 0.34 s, p99 0.75 sThird-party telemetryOpenRouter, as of Oct 1 (page showed 0.18 s p50 on Oct 4)Oct 4, 2026
Jev 193.6x faster and 444.6x cheaper than frontier LLMsVendor claimTypeSafe workflow evalsOct 4, 2026
Jev 100% agreement with human labels on 500 decisionsThird-party study, single taskLangChainOct 4, 2026

Read the table this way. On Jev you can build and test today against a documented contract. On the Decisions API you can plan, but you cannot design thresholds, write a parser or forecast cost, because the parts that matter for a voice agent are not public.

There is also a naming trap. OpenRouter runs its own "Decisions" endpoint (POST openrouter.ai/api/alpha/decisions) that serves Jev. If a teammate says "we're on the Decisions API," ask which one.

Architecture: constrained generative model vs purpose-built decision model

The two products answer the same kind of question in different ways. That difference drives off-schema risk, cost, prose capability and calibration.

Side-by-side diagram comparing a constrained generative LLM that decodes answer tokens under a grammar with a purpose-built decision model that scores every option and returns a probability distribution

How a constrained generative model decides

OpenAI describes the Decisions API as focusing Luna's intelligence on questions with finite answers. OpenAI has not published the mechanism. The common way to make a generative model pick from a list is constrained decoding. The model generates tokens as usual, but a grammar or mask blocks any token that would leave the allowed set. When the answer label is complete, generation stops.

That design has three consequences you can reason about even without docs:

1. Off-schema output becomes rare at the API boundary. If the constraint is enforced in the decoder, the returned label is always one of your options. Your parser stops failing. That does not mean the label is right.

2. The model still spends output tokens. Even a one-word label is several tokens. If any hidden reasoning runs first, those tokens are billed as output on GPT-6 Luna, which supports reasoning effort settings. Whether the Decisions API exposes or bills reasoning is not published.

3. Forcing a format can change the answer. Tam et al. (2024), "Let Me Speak Freely?" tested LLMs with strict format restrictions against free-form answers and found a significant decline in reasoning performance, with stricter constraints causing larger drops. A purpose-built product may avoid this, but you should test it, not assume it.

Probabilities are a separate question. A generative model has token probabilities internally. Turning those into a clean per-option distribution means scoring every option, not just decoding the winner. OpenAI has not said whether the Decisions API does this.

How a purpose-built decision model decides

Jev is a System One model. You send `state` (a string, JSON object or array) and a set of `questions`. Each question is a Choice (pick from named options), a Score (position on an ordered scale) or a Noul (probability a condition holds). The response for a Choice includes `choice`, `confidence` and a `probabilities` map over every option.

TypeSafe says Jev is trained with RLCD (Reinforcement Learning for Calibrated Decisions). It does not write prose. Output tokens are free because the product is priced on input. Several questions can ride in one call, so a single request can return a route, an urgency score and a needs-human flag together.

The trade-offs follow from that:

PropertyConstrained generative LLM (Decisions API, inferred)Purpose-built decision model (Jev, documented)
Off-schema risk at the boundaryLow if decoder-constrained; contract unpublishedNone by design: answers are typed objects
Output-token costLikely nonzero; unpublishedFree
Prose or rationalePossible in principle; unpublishedNone
Per-option probabilitiesUnpublishedReturned for every option
Calibration claimNone publishedVendor trains for calibration (RLCD)
Image inputYesNo, text only
Known failure modeFormat constraints can cut reasoning qualityAccuracy falls as irrelevant state grows (TypeSafe jaggedness page)

The last row matters for voice. A full call transcript is mostly irrelevant to any single decision. TypeSafe's own jaggedness page warns that Jev's accuracy drops as irrelevant state grows. The practical fix is to send the last few turns plus the structured fields the decision needs, not the whole call. Expect the same discipline to help any decision model, including Luna-based ones.

Latency inside a voice turn

On a phone call, the number that matters is not the decision call in isolation. It is how much the decision adds to the gap between the caller finishing a sentence and hearing the agent start.

Human conversation sets a hard bar. Stivers et al. (2009, PNAS) measured turn transitions across ten languages and found gaps cluster around 200 ms, with the same general pattern in every language studied. Voice agents cannot hit that, but the further you drift past about 800 ms, the more callers talk over the agent or ask "hello?"

A worked turn budget

Take an illustrative cascaded pipeline. These are assumptions, not measurements:

HopAssumed ms
Endpointing (silence or end-of-turn model)200
STT final transcript80
LLM time to first token300
TTS time to first byte120
Transport and telephony (SIP, RTP, jitter buffer)100
Total, no decision call800

Now add a blocking decision before the LLM runs, for example an escalation check or a route pick.

  • Jev at OpenRouter p50 (210 ms): 800 + 210 = 1,010 ms.
  • Jev at OpenRouter p95 (340 ms): 800 + 340 = 1,140 ms.
  • Jev at OpenRouter p99 (750 ms): 800 + 750 = 1,550 ms.
  • Decisions API at OpenAI's 150 ms claim: 800 + 150 = 950 ms. The conditions behind 150 ms are not stated. It is not labeled p50 or p95, and the baseline is OpenAI's own Luna API.

The 60 ms gap between 150 ms and 210 ms looks decisive. It is not. One number is a vendor slide with unknown conditions. The other is a router's round-trip telemetry, which includes OpenRouter's own hop and depends on where requests originate. Neither is your network path from your agent servers. Measure from your own region.

Waterfall chart of an 800 millisecond voice turn showing a decision call added in series at Jev p50, Jev p95 and the OpenAI 150 millisecond claim, compared with running the decision in parallel with STT

Why p95 matters more than p50 on calls

A call is not one decision. It is many turns, and the caller remembers the worst pause. If a blocking decision runs on every turn, the chance that a call sees at least one tail-latency turn is:

P(at least one slow turn) = 1 - (1 - p_tail) ^ turns

For a 10-turn call:

  • Chance of at least one turn at or beyond p95: 1 - 0.95^10 = 40.1%.
  • Chance of at least one turn at or beyond p99: 1 - 0.99^10 = 9.6%.

So with Jev's OpenRouter p99 of 0.75 s, roughly one call in ten would hit at least one turn that adds three-quarters of a second, under these assumptions. For the Decisions API there is no published p95 or p99 at all. That is the single biggest unknown for voice use, more than the headline number.

Keep decisions off the critical path where you can

Three patterns cut the added latency close to zero:

1. Run in parallel with STT finalization. Fire the decision on the interim transcript when the end-of-turn signal starts. If the final transcript matches, use the result. If not, re-run. In the waterfall above, this hides most of a 210 ms call behind endpointing.

2. Run in parallel with the LLM. Start the LLM response and the escalation check together. If the check says escalate, cancel the LLM stream before TTS speaks. You pay a little wasted LLM compute to save the serial wait.

3. Move it after the call. Disposition codes, compliance flags and QA scores rarely need to block a turn. Run them post-call, where latency only affects how fast your dashboard updates.

Always set a hard timeout on a blocking decision. If it returns late, take a safe default (continue the flow, or ask a clarifying question) and log the miss. A decision model that times out should never leave the caller in silence.

Cost at 50k to 500k calls a month

Cost is easier to estimate than latency, but for one of the two products the price does not exist yet. The math below uses Jev's real price and GPT-6 Luna's list price as a yardstick only. Luna list price is not Decisions API pricing. OpenAI may price Decisions per call, per token, or with different rates. Treat the Luna rows as "what it might look like if Decisions billed like Luna," nothing more.

Labeled assumptions

  • Decisions per call: 6 (one route, three tool gates, one escalation check, one post-call disposition).
  • Input tokens per decision: 1,500 (recent turns, a few structured fields, option descriptions).
  • Output tokens per decision on the Luna yardstick: two cases, 20 (label only) and 200 (if hidden reasoning were billed).
  • No caching on either side. Luna lists cached input at $0.01 / 1M, so a stable option-list prefix could cut the Luna input line if Decisions supports caching. That is unpublished.

Input tokens per call = 6 x 1,500 = 9,000.

The arithmetic

Jev per call: 9,000 x $0.042 / 1,000,000 = $0.000378.

Luna yardstick per call, 20 output tokens per decision: input 9,000 x $0.10 / 1M = $0.00090, output 120 x $0.50 / 1M = $0.00006, total $0.00096.

Luna yardstick per call, 200 output tokens per decision: input $0.00090, output 1,200 x $0.50 / 1M = $0.00060, total $0.00150.

Calls per monthJev (real price)Luna yardstick, 20 outLuna yardstick, 200 out
50,000$18.90$48.00$75.00
150,000$56.70$144.00$225.00
500,000$189.00$480.00$750.00

What the numbers say

The ratio looks large: Jev is 2.5 to 4 times cheaper than the Luna yardstick here. The absolute numbers are small. Put them next to the voice minutes. OpenAI's GPT-Live lists at $0.05 per minute. A 4-minute call is $0.20 in speech model cost alone, before telephony. The decision layer at $0.0004 to $0.0015 per call is under 1% of that.

So at ICP3 volumes, price per decision is rarely the deciding factor. What decides is whether the decision is right, whether it arrives inside your turn budget at p95, and whether it gives you a probability you can threshold. If OpenAI prices Decisions far above Luna, or per call, rerun this table. The formula is:

monthly cost = calls x decisions_per_call x (input_tokens x input_rate + output_tokens x output_rate) / 1,000,000

Where image input matters for voice, and where it does not

Image input is the Decisions API's clearest confirmed advantage. Jev is text-only. For a pure phone line, this does not matter: a SIP call carries audio, and your decision model sees text from STT.

Images enter voice workflows in a few specific places:

  • Screen-share or co-browse support. A web voice widget where the caller shares a screen. "Is the user on the billing settings page?" is a bounded question about an image.
  • Document photo verification mid-call. The agent texts a link, the caller uploads a photo of an insurance card or ID, and the agent decides "legible / not legible / wrong document."
  • Claims and field service. A photo of damage, a meter reading or a part label, classified into a fixed set before the agent continues.
  • Receipt or order disputes. A photo of a receipt checked against "matches order / does not match / unreadable."

If none of these are on your roadmap, image input is not a reason to wait. If they are, you can run them on Jev today only by adding a vision model first to extract text, then deciding on the text. That adds a hop and its latency. A single Decisions call that takes the image directly could be simpler, once its contract and latency on image input are public. Image tokens cost more and may be slower, so do not reuse the 150 ms text figure for image decisions.

Probabilities, thresholds and abstention

This is the section most comparisons skip, and it matters most for voice.

A voice agent rarely wants just "the answer." It wants "the answer, and how sure." Escalation policy depends on it. If the agent is 97% sure the caller wants to cancel, it can route straight to the retention flow. If it is 55% sure, it should ask. If two routes are within a few points of each other, it should confirm. Without a probability, every decision is treated as equally certain, and you lose the ability to abstain.

The math of abstention

Geifman and El-Yaniv (2017), "Selective Classification for Deep Neural Networks" formalized this. Given a confidence score, you choose a threshold so the model answers only when confident, trading coverage (how often it answers) for risk (how often it is wrong when it does). On your own labeled data, you sweep the threshold and pick the point where error on answered cases meets your target.

For a voice agent, that becomes a three-band policy:

BandRule (illustrative)Agent action
Acttop probability at least 0.90 and margin over runner-up at least 0.20Route or run the tool
Confirmtop probability 0.60 to 0.90, or margin under 0.20Ask a short confirming question
Escalate or abstaintop probability under 0.60, or needs-human probability above your floorHand off or fall back

Tune these on your own calls. Set tighter bands for high-cost routes (cancellations, payments) and looser bands for low-cost ones (store hours). Our routing and escalation guide walks through per-route thresholds, and escalation accuracy covers how to measure them.

Calibration: is 0.9 really 90%?

A probability is only useful if it is calibrated. Guo et al. (2017), "On Calibration of Modern Neural Networks" showed that modern deep networks are often overconfident, and that a simple post-hoc fix (temperature scaling) recovers much of the gap. The test is a reliability diagram: bucket predictions by stated confidence and check that the 0.9 bucket is right about 90% of the time. Expected Calibration Error (ECE) summarizes the gap as one number.

You do not need logits to check this. With returned probabilities and 500 to 1,000 labeled decisions, you can plot reliability yourself and, if needed, fit an isotonic mapping on top.

Option order is a real risk for any decision model

On Hacker News, a commenter (hbrn) reported that Jev's probabilities shift when the answer list is reordered, and that its confidence is derived from those probabilities. That critique is fair, and it is not unique to Jev. Research on LLMs shows the same effect:

  • Pezeshkpour and Hruschka (2023) found performance gaps of roughly 13% to 75% across benchmarks when multiple-choice options were reordered. They traced it to positional bias when the model is torn between its top two or three choices, and got up to 8 points back with calibration methods.
  • Zheng et al. (ICLR 2024) showed LLMs carry a prior toward specific option ID tokens and proposed PriDe, which estimates that prior by permuting options on a small sample and then removes it.

The practical rule for both products: fix your option order in code, never let it vary per request, and include an order-permutation test in your evaluation. If probabilities move a lot under permutation, average over two or three fixed orders for high-stakes decisions, or widen your Confirm band.

What to do if the Decisions API returns only a label

If OpenAI ships without probabilities, you have three workarounds, each with a cost:

1. Add an explicit "unsure" option. Cheap, but the model's use of "unsure" is itself uncalibrated.

2. Ask k times with permuted option order and treat agreement rate as confidence. This multiplies cost and, if run in series, latency.

3. Use it only where you do not need a threshold, such as post-call tagging reviewed in aggregate.

None of these are as good as a returned distribution for live escalation. If the Decisions API ships with probabilities, the gap closes, and you should rerun the calibration test on it.

Choosing between decision models for your voice agent?
Evalgent runs both on your own labeled calls and reports accuracy, calibration and p95 latency inside your turn budget, independent of either vendor.
Book a demo

Voice use-case map: which fits today

Not every decision has the same needs. The table maps the five most common voice decisions against what each product offers today.

Use caseWhen it runsLatency needNeeds probabilitiesImagesFit today
Routing and escalationEvery turn or first turnsTight, blockingYesNoJev, documented and testable; Decisions after its p95 and probability support are published
Tool-call gatingBefore each tool callTight, blockingYes, for confirm bandNoJev; see tool calling with Jev
Disposition codesPost-callLooseHelpful for review queuesRarelyEither, once Decisions has pricing; Jev today
Compliance flags (disclosures, verification)Live or post-callModerateYes (Noul-style probability)NoJev today; audit trail of probabilities helps
Live QA and call scoringPost-call or near-liveLooseYes (score distributions)RarelyJev today; see scoring every call
Photo or screen checks in a voice flowMid-callModerateHelpfulYesDecisions API, once available to you; Jev needs a vision pre-step

Two patterns stand out. First, the live, blocking decisions are where probabilities and tail latency matter most, and those are exactly the two things the Decisions API has not published. Second, post-call work is forgiving. If you already run on OpenAI and only need dispositions, waiting a week for Decisions docs costs you little.

For deeper treatment of the Jev side, see Jev for voice agents, Jev vs LLMs and call monitoring with Jev.

OpenAI's real advantages

A neutral comparison has to give OpenAI its due. Several advantages are real even before docs ship:

  • Image input. Confirmed, and Jev does not have it.
  • One vendor, one bill, one security review. If your voice stack already runs on OpenAI (GPT-Live, Realtime, or a GPT-6 model for the main LLM), adding Decisions avoids a new vendor contract, a new DPA and a new data-flow diagram. For a 50 to 500 person company with a lean security team, that can outweigh a price difference that is under 1% of call cost.
  • Ecosystem and tooling. OpenAI's SDKs, dashboards, usage tracking, data residency options and enterprise tiers already exist. Decisions will likely plug into them.
  • Model lineage. GPT-6 Luna is a general model with broad world knowledge. For decisions that need that knowledge, a Luna-based decider may do better on questions that are vague or depend on context outside the transcript. That is a hypothesis to test, not a fact.

Jev's advantages are equally concrete: it is available, documented, priced, returns probabilities for every option, is free on output, and has third-party latency telemetry. The LangChain study reported 100% agreement with human labels over 500 decisions, versus 99.8% for GPT-5.6 Terra, at $0.34 versus $28.17. That is one task, run by a partner, and it is not voice data. Use it as a reason to test, not as a result for your calls.

Reactions so far are mostly about what is missing. Firecrawl's analysis and Vercel's note both flag that OpenAI has not specified probability output. None of these are tests of the product.

Decision tree for choosing between Jev and the OpenAI Decisions API today based on whether you need images, live blocking decisions, per-option probabilities, and single-vendor procurement

Decision matrix: what to choose this week

Here is a reusable matrix. Find your row; the right column is the call we would make with today's public information.

Your situationChoose nowWhy
Live routing or escalation on a phone line, need to ship this monthJev, behind an adapterDocumented contract, probabilities, measured tail latency
Tool-call gating on payments or account changesJev, behind an adapterConfirm band needs probabilities
Post-call dispositions only, stack already all-OpenAIWait for Decisions docs, then test bothLow latency need, procurement simplicity
Need image decisions (screen-share, document photo)Plan for Decisions; prototype with a vision model plus JevOnly Decisions takes images directly
Strict "single AI vendor" security policyDecisions, once GA and testedAvoids a new vendor review
Strict data-residency needCheck both vendors' current termsNeither product's residency terms for this endpoint are covered here
You cannot get preview accessJev now, Decisions as a planned challengerYou cannot test what you cannot call
High-stakes compliance flags with audit needsJev now; retest when Decisions publishes probabilitiesLogged probabilities make audits defensible

The principle behind every row: commit to an interface, not a vendor. The next section shows how.

A provider-agnostic decision adapter in Python

The cheapest insurance against an unfinished API is a thin adapter. Your agent code calls one function with a provider-neutral question. Each provider has a small class that translates it. When OpenAI publishes the Decisions contract, you fill in one class and run both through the same evaluation.

The Jev call below follows the TypeSafe quickstart: `TypeSafeClient`, `client.system_one(state=..., questions=...)`, `Choice(instructions=..., criteria=...)`, and `response.answers[name]` with `.choice`, `.confidence` and `.probabilities`. The OpenAI class is a placeholder on purpose. Its schema is unpublished, so we do not guess it.

# Simplified, illustrative adapter. Requires: pip install typesafe-sdk
# The client reads TYPESAFE_API_KEY from the environment.
import asyncio
import time
from dataclasses import dataclass, field
from typing import Optional, Protocol

from typesafe_sdk import Choice, TypeSafeClient


@dataclass(frozen=True)
class Question:
    name: str
    instructions: str
    options: dict[str, str]  # label -> description; keep order fixed in code


@dataclass
class Decision:
    name: str
    choice: str
    probabilities: Optional[dict[str, float]] = None  # None if provider lacks them
    confidence: Optional[float] = None
    provider: str = ""
    latency_ms: float = 0.0
    timed_out: bool = False


class DecisionProvider(Protocol):
    name: str

    def decide(self, state: str, questions: list[Question]) -> list[Decision]: ...


class JevProvider:
    name = "jev"

    def __init__(self) -> None:
        self.client = TypeSafeClient()

    def decide(self, state: str, questions: list[Question]) -> list[Decision]:
        start = time.perf_counter()
        response = self.client.system_one(
            state=state,
            questions={
                q.name: Choice(instructions=q.instructions, criteria=q.options)
                for q in questions
            },
        )
        elapsed = (time.perf_counter() - start) * 1000
        out = []
        for q in questions:
            a = response.answers[q.name]
            out.append(Decision(
                name=q.name,
                choice=a.choice,
                probabilities=dict(a.probabilities),
                confidence=a.confidence,
                provider=self.name,
                latency_ms=elapsed,
            ))
        return out


class OpenAIDecisionsProvider:
    name = "openai-decisions"

    def decide(self, state: str, questions: list[Question]) -> list[Decision]:
        # TODO: fill in when OpenAI publishes the Decisions API contract
        # (endpoint, request schema, response schema, SDK method).
        # As of October 4, 2026 none of these are public. Do not guess them.
        # Map the response to Decision objects. If no per-option probabilities
        # are returned, leave probabilities=None so the policy below abstains
        # from auto-acting.
        raise NotImplementedError("OpenAI Decisions API contract not yet published")


async def decide_with_budget(
    provider: DecisionProvider,
    state: str,
    questions: list[Question],
    budget_ms: int,
    fallback: dict[str, str],
) -> list[Decision]:
    """Run a blocking decision with a hard timeout so the caller never hears silence."""
    try:
        return await asyncio.wait_for(
            asyncio.to_thread(provider.decide, state, questions),
            timeout=budget_ms / 1000,
        )
    except (asyncio.TimeoutError, NotImplementedError):
        return [
            Decision(name=q.name, choice=fallback[q.name],
                     provider=provider.name, timed_out=True)
            for q in questions
        ]


def policy(d: Decision, act_p: float = 0.90, min_margin: float = 0.20,
           floor_p: float = 0.60) -> str:
    """Three-band policy. Illustrative thresholds; tune on your own labeled calls."""
    if d.timed_out or d.probabilities is None:
        return "confirm"
    ranked = sorted(d.probabilities.values(), reverse=True)
    top = ranked[0]
    margin = top - (ranked[1] if len(ranked) > 1 else 0.0)
    if top >= act_p and margin >= min_margin:
        return "act"
    if top < floor_p:
        return "escalate"
    return "confirm"


ROUTE = Question(
    name="route",
    instructions="Pick the flow that best serves the caller's main need on this turn.",
    options={
        "billing": "Charges, refunds, invoices or payment methods.",
        "cancel": "Caller wants to end or downgrade a plan.",
        "schedule": "Book, move or cancel an appointment.",
        "support": "Something is broken: device, app or login.",
        "other": "Anything else.",
    },
)


async def on_user_turn(recent_turns: str, provider: DecisionProvider) -> str:
    [route] = await decide_with_budget(
        provider, recent_turns, [ROUTE], budget_ms=300,
        fallback={"route": "other"},
    )
    action = policy(route)
    # Log provider, choice, probabilities, latency_ms and action for every turn.
    return f"{action}:{route.choice}"

A few design notes that save pain later:

  • Option order is fixed in code. The `options` dict preserves insertion order, which protects you from the order effects described above.
  • The policy degrades safely. A provider that returns no probabilities, or a call that times out, lands in "confirm," never "act."
  • Latency is measured at your boundary. `latency_ms` is what your agent experienced, which is the number you need, not the vendor's.
  • The state is short. Pass recent turns and structured fields, not the full transcript, to stay clear of the irrelevant-state accuracy drop.

In a LiveKit agent server or a Pipecat pipeline, call `on_user_turn` from wherever you already handle the final user transcript, and start it as early as your end-of-turn signal allows.

How to test the OpenAI Decisions API vs Jev on your own calls

Vendor numbers will not tell you which product is right for your calls. A one-week, side-by-side test will. Our sibling post on benchmarking decision APIs for voice agents goes deeper on harness design; this is the short version.

1. Pick three to five decisions you actually make. For example: route, needs-human, tool gate for refunds, and post-call disposition. Write each as a fixed-order option list in the adapter format above.

2. Build a labeled set from real calls. Pull 300 to 1,000 decision points from recorded calls, stratified by route and including hard cases (mixed intents, angry callers, noisy transcripts). Have two people label each, and resolve disagreements. Our golden dataset guide covers sampling and labeling.

3. Freeze the state window. Decide exactly what each decision sees (for example, the last four turns plus account status) and use the same input for both providers.

4. Run Jev now, and Decisions as soon as you have access. Send identical inputs through the adapter. Record choice, probabilities (if any), and latency measured from your agent server's region.

5. Measure accuracy and calibration. Report accuracy per decision and per route. For any provider that returns probabilities, plot a reliability diagram and compute ECE. Check the Act band's error rate against your target.

6. Run the order test. Re-run 100 items with options in two other fixed orders. Count how often the chosen label changes and how far top probability moves.

7. Measure latency where it hurts. Report p50, p95 and p99 over at least 1,000 calls per provider, at your real concurrency, from your production region. Plug p95 and p99 into the turn budget and the at-least-one-slow-turn formula.

8. Price it on your traffic. Use your actual tokens per decision and decisions per call in the cost formula. Replace the Luna yardstick with real Decisions pricing the day it is published.

9. Decide with a written rule. For example: choose the provider with higher accuracy on high-stakes routes, provided its p95 stays under 300 ms and its Act-band error is under 2%. Write the rule before you look at results.

If you want this done independently, this is the kind of bake-off Evalgent runs: same calls, same inputs, both providers, with accuracy, calibration and tail latency reported against your own thresholds. The LLM-as-judge limits post explains why we do not score decision models with another LLM's opinion, and Jev as a judge covers when a decision model can itself be the scorer.

Frequently asked questions

Is the OpenAI Decisions API available to everyone?

Not as of October 4, 2026. OpenAI announced it at DevDay on September 29 as a limited preview for selected API customers, with broad release planned "in the coming days." There is no public docs page, schema or price yet. A third party, eesel, reported that a standard key received a 403 "not enabled" response.

How much does the OpenAI Decisions API cost?

OpenAI has not published Decisions API pricing. It is not on the API pricing page or the docs pricing table as of October 4, 2026. GPT-6 Luna lists at $0.10 per million input tokens and $0.50 per million output tokens, but that is a yardstick, not the Decisions price.

Does the Decisions API return probabilities or confidence scores?

OpenAI has not said. The DevDay text describes answers to questions with finite predefined answers, without mentioning probabilities. Firecrawl and Vercel both flagged this gap. Jev returns a probability for every option plus a confidence value, which is what escalation thresholds and abstention need.

Is 150 ms faster than Jev?

Not necessarily. The 150 ms figure is a DevDay slide comparing Decisions with OpenAI's own Luna API, with unstated conditions and no percentile. Jev's 0.21 s p50 and 0.34 s p95 are OpenRouter round-trip telemetry as of October 1. Measure both from your own servers before comparing.

Can Jev handle images?

No. Jev input is text only: a string, JSON object or array. For voice flows that involve a document photo or screen-share, you would run a vision model first to extract text, then decide with Jev. The Decisions API accepts images directly, which is its clearest confirmed advantage.

How consistent are Jev's answers?

Jev is highly consistent, though not perfectly fixed: the LangChain study found LLM judges' score variance 92 to 913 times higher than Jev's on its task. A Hacker News commenter showed probabilities shift when options are reordered, so fix option order in code and test permutations.

Which should I use for voice agent escalation today?

If you need to ship live escalation now, Jev is the one you can test against a documented contract, with per-option probabilities for a confirm band. Put it behind an adapter so you can trial the Decisions API when its docs, probability support and tail latency are published.

Will the cost difference matter at my call volume?

Probably not much. Under the assumptions in this post, Jev costs about $189 a month at 500,000 calls, and a Luna-priced equivalent would be $480 to $750. Both are under 1% of typical speech model cost per call. Accuracy, p95 latency and probabilities matter more.

The bottom line

Today, Jev is the decision model you can document, price, test and threshold for live voice decisions, while the OpenAI Decisions API is a credible challenger whose schema, probabilities, tail latency and price are still unpublished. Build behind a provider-agnostic adapter, test both on your own labeled calls when OpenAI opens access, and let accuracy, calibration and p95 latency inside your turn budget make the call.

Related Articles