Open door for builders.
OpenAI Decisions API vs Jev for Voice Agents: What's Known, What Isn't, and How to Choose

On this page
Last verified October 4, 2026. OpenAI says broad release is coming "in the coming days." This post will be updated when OpenAI publishes the Decisions API docs and pricing.
Every voice agent makes small, bounded decisions on almost every turn. Which flow does this caller need? Is it safe to run the refund tool? Should a human take over now? What disposition code goes on this call? A general LLM can answer all of these, but it answers slowly and in prose you then have to parse.
Two products now target exactly that slice. TypeSafe launched Jev, a purpose-built decision model, on September 15, 2026. OpenAI announced the Decisions API, powered by a constrained version of GPT-6 Luna, at DevDay on September 29, 2026.
If you run a voice agent in-house on LiveKit, Pipecat or Dograh, you now have to decide whether to wire one in, wait for the other, or build so you can swap. This guide is written from the position of an independent evaluator. Evalgent does not resell either product. The goal is to separate what is confirmed from what is marketing, run the numbers inside a real voice-turn budget, and give you a way to choose this week without regretting it next month.
What is confirmed, what is a vendor claim, and what is not published
Start with the evidence. The table below marks every claim you are likely to see in a comparison post. "Confirmed" means a primary source states it as fact. "Vendor claim" means the vendor said it but conditions are not public or no third party has reproduced it. "Not published" means nobody outside the preview can check it yet.
| Claim | Status | Source | Date checked |
|---|---|---|---|
| Decisions API exists; answers user-defined questions with finite predefined answers | Confirmed | OpenAI DevDay 2026 recap | Oct 4, 2026 |
| Decisions API is powered by GPT-6 Luna | Confirmed | OpenAI DevDay recap | Oct 4, 2026 |
| Decisions API accepts text or image context | Confirmed | OpenAI DevDay recap | Oct 4, 2026 |
| Limited preview; broad release "in the coming days" | Confirmed | OpenAI DevDay recap; @OpenAIDevs | Oct 4, 2026 |
| 150 ms vs 1.6 s for the GPT-6 Luna API ("10x faster") | Vendor claim | DevDay slide reproduced by The Decoder | Oct 4, 2026 |
| Endpoint path, request and response schema, SDK method | Not published | No docs page; guide URL returns no content | Oct 4, 2026 |
| Per-option probabilities or confidence returned | Not published | Noted as unspecified by Vercel and Firecrawl | Oct 4, 2026 |
| Multiple questions per call, context limit | Not published | Firecrawl docs-index check, Sep 30 | Oct 4, 2026 |
| Decisions API pricing | Not published | No row on OpenAI API pricing or the docs pricing table | Oct 4, 2026 |
| Standard API key gets 403 "not enabled for this user" on POST /v1/decisions | Third-party report | eesel (not our test) | Oct 4, 2026 |
| GPT-6 Luna list price $0.10 / 1M input, $0.50 / 1M output | Confirmed (Luna, not Decisions) | GPT-6 Luna model page | Oct 4, 2026 |
| Jev endpoint POST api.typesafe.ai/v1/systemone, model `jev-latest` | Confirmed | TypeSafe quickstart | Oct 4, 2026 |
| Jev returns choice, confidence and per-option probabilities | Confirmed | TypeSafe quickstart response body | Oct 4, 2026 |
| Jev price $0.042 / 1M input tokens, output free | Confirmed | OpenRouter Jev 1.13 | Oct 4, 2026 |
| Jev input is text only | Confirmed | TypeSafe docs | Oct 4, 2026 |
| Jev context 32K on OpenRouter, 64K per TypeSafe docs | Confirmed (two figures) | OpenRouter; TypeSafe docs | Oct 4, 2026 |
| Jev latency 70 to 500 ms | Vendor claim | TypeSafe | Oct 4, 2026 |
| Jev p50 0.21 s, p95 0.34 s, p99 0.75 s | Third-party telemetry | OpenRouter, as of Oct 1 (page showed 0.18 s p50 on Oct 4) | Oct 4, 2026 |
| Jev 193.6x faster and 444.6x cheaper than frontier LLMs | Vendor claim | TypeSafe workflow evals | Oct 4, 2026 |
| Jev 100% agreement with human labels on 500 decisions | Third-party study, single task | LangChain | Oct 4, 2026 |
Read the table this way. On Jev you can build and test today against a documented contract. On the Decisions API you can plan, but you cannot design thresholds, write a parser or forecast cost, because the parts that matter for a voice agent are not public.
There is also a naming trap. OpenRouter runs its own "Decisions" endpoint (POST openrouter.ai/api/alpha/decisions) that serves Jev. If a teammate says "we're on the Decisions API," ask which one.
Architecture: constrained generative model vs purpose-built decision model
The two products answer the same kind of question in different ways. That difference drives off-schema risk, cost, prose capability and calibration.

How a constrained generative model decides
OpenAI describes the Decisions API as focusing Luna's intelligence on questions with finite answers. OpenAI has not published the mechanism. The common way to make a generative model pick from a list is constrained decoding. The model generates tokens as usual, but a grammar or mask blocks any token that would leave the allowed set. When the answer label is complete, generation stops.
That design has three consequences you can reason about even without docs:
1. Off-schema output becomes rare at the API boundary. If the constraint is enforced in the decoder, the returned label is always one of your options. Your parser stops failing. That does not mean the label is right.
2. The model still spends output tokens. Even a one-word label is several tokens. If any hidden reasoning runs first, those tokens are billed as output on GPT-6 Luna, which supports reasoning effort settings. Whether the Decisions API exposes or bills reasoning is not published.
3. Forcing a format can change the answer. Tam et al. (2024), "Let Me Speak Freely?" tested LLMs with strict format restrictions against free-form answers and found a significant decline in reasoning performance, with stricter constraints causing larger drops. A purpose-built product may avoid this, but you should test it, not assume it.
Probabilities are a separate question. A generative model has token probabilities internally. Turning those into a clean per-option distribution means scoring every option, not just decoding the winner. OpenAI has not said whether the Decisions API does this.
How a purpose-built decision model decides
Jev is a System One model. You send `state` (a string, JSON object or array) and a set of `questions`. Each question is a Choice (pick from named options), a Score (position on an ordered scale) or a Noul (probability a condition holds). The response for a Choice includes `choice`, `confidence` and a `probabilities` map over every option.
TypeSafe says Jev is trained with RLCD (Reinforcement Learning for Calibrated Decisions). It does not write prose. Output tokens are free because the product is priced on input. Several questions can ride in one call, so a single request can return a route, an urgency score and a needs-human flag together.
The trade-offs follow from that:
| Property | Constrained generative LLM (Decisions API, inferred) | Purpose-built decision model (Jev, documented) |
|---|---|---|
| Off-schema risk at the boundary | Low if decoder-constrained; contract unpublished | None by design: answers are typed objects |
| Output-token cost | Likely nonzero; unpublished | Free |
| Prose or rationale | Possible in principle; unpublished | None |
| Per-option probabilities | Unpublished | Returned for every option |
| Calibration claim | None published | Vendor trains for calibration (RLCD) |
| Image input | Yes | No, text only |
| Known failure mode | Format constraints can cut reasoning quality | Accuracy falls as irrelevant state grows (TypeSafe jaggedness page) |
The last row matters for voice. A full call transcript is mostly irrelevant to any single decision. TypeSafe's own jaggedness page warns that Jev's accuracy drops as irrelevant state grows. The practical fix is to send the last few turns plus the structured fields the decision needs, not the whole call. Expect the same discipline to help any decision model, including Luna-based ones.
Latency inside a voice turn
On a phone call, the number that matters is not the decision call in isolation. It is how much the decision adds to the gap between the caller finishing a sentence and hearing the agent start.
Human conversation sets a hard bar. Stivers et al. (2009, PNAS) measured turn transitions across ten languages and found gaps cluster around 200 ms, with the same general pattern in every language studied. Voice agents cannot hit that, but the further you drift past about 800 ms, the more callers talk over the agent or ask "hello?"
A worked turn budget
Take an illustrative cascaded pipeline. These are assumptions, not measurements:
| Hop | Assumed ms |
|---|---|
| Endpointing (silence or end-of-turn model) | 200 |
| STT final transcript | 80 |
| LLM time to first token | 300 |
| TTS time to first byte | 120 |
| Transport and telephony (SIP, RTP, jitter buffer) | 100 |
| Total, no decision call | 800 |
Now add a blocking decision before the LLM runs, for example an escalation check or a route pick.
- Jev at OpenRouter p50 (210 ms): 800 + 210 = 1,010 ms.
- Jev at OpenRouter p95 (340 ms): 800 + 340 = 1,140 ms.
- Jev at OpenRouter p99 (750 ms): 800 + 750 = 1,550 ms.
- Decisions API at OpenAI's 150 ms claim: 800 + 150 = 950 ms. The conditions behind 150 ms are not stated. It is not labeled p50 or p95, and the baseline is OpenAI's own Luna API.
The 60 ms gap between 150 ms and 210 ms looks decisive. It is not. One number is a vendor slide with unknown conditions. The other is a router's round-trip telemetry, which includes OpenRouter's own hop and depends on where requests originate. Neither is your network path from your agent servers. Measure from your own region.

Why p95 matters more than p50 on calls
A call is not one decision. It is many turns, and the caller remembers the worst pause. If a blocking decision runs on every turn, the chance that a call sees at least one tail-latency turn is:
P(at least one slow turn) = 1 - (1 - p_tail) ^ turns
For a 10-turn call:
- Chance of at least one turn at or beyond p95: 1 - 0.95^10 = 40.1%.
- Chance of at least one turn at or beyond p99: 1 - 0.99^10 = 9.6%.
So with Jev's OpenRouter p99 of 0.75 s, roughly one call in ten would hit at least one turn that adds three-quarters of a second, under these assumptions. For the Decisions API there is no published p95 or p99 at all. That is the single biggest unknown for voice use, more than the headline number.
Keep decisions off the critical path where you can
Three patterns cut the added latency close to zero:
1. Run in parallel with STT finalization. Fire the decision on the interim transcript when the end-of-turn signal starts. If the final transcript matches, use the result. If not, re-run. In the waterfall above, this hides most of a 210 ms call behind endpointing.
2. Run in parallel with the LLM. Start the LLM response and the escalation check together. If the check says escalate, cancel the LLM stream before TTS speaks. You pay a little wasted LLM compute to save the serial wait.
3. Move it after the call. Disposition codes, compliance flags and QA scores rarely need to block a turn. Run them post-call, where latency only affects how fast your dashboard updates.
Always set a hard timeout on a blocking decision. If it returns late, take a safe default (continue the flow, or ask a clarifying question) and log the miss. A decision model that times out should never leave the caller in silence.
Cost at 50k to 500k calls a month
Cost is easier to estimate than latency, but for one of the two products the price does not exist yet. The math below uses Jev's real price and GPT-6 Luna's list price as a yardstick only. Luna list price is not Decisions API pricing. OpenAI may price Decisions per call, per token, or with different rates. Treat the Luna rows as "what it might look like if Decisions billed like Luna," nothing more.
Labeled assumptions
- Decisions per call: 6 (one route, three tool gates, one escalation check, one post-call disposition).
- Input tokens per decision: 1,500 (recent turns, a few structured fields, option descriptions).
- Output tokens per decision on the Luna yardstick: two cases, 20 (label only) and 200 (if hidden reasoning were billed).
- No caching on either side. Luna lists cached input at $0.01 / 1M, so a stable option-list prefix could cut the Luna input line if Decisions supports caching. That is unpublished.
Input tokens per call = 6 x 1,500 = 9,000.
The arithmetic
Jev per call: 9,000 x $0.042 / 1,000,000 = $0.000378.
Luna yardstick per call, 20 output tokens per decision: input 9,000 x $0.10 / 1M = $0.00090, output 120 x $0.50 / 1M = $0.00006, total $0.00096.
Luna yardstick per call, 200 output tokens per decision: input $0.00090, output 1,200 x $0.50 / 1M = $0.00060, total $0.00150.
| Calls per month | Jev (real price) | Luna yardstick, 20 out | Luna yardstick, 200 out |
|---|---|---|---|
| 50,000 | $18.90 | $48.00 | $75.00 |
| 150,000 | $56.70 | $144.00 | $225.00 |
| 500,000 | $189.00 | $480.00 | $750.00 |
What the numbers say
The ratio looks large: Jev is 2.5 to 4 times cheaper than the Luna yardstick here. The absolute numbers are small. Put them next to the voice minutes. OpenAI's GPT-Live lists at $0.05 per minute. A 4-minute call is $0.20 in speech model cost alone, before telephony. The decision layer at $0.0004 to $0.0015 per call is under 1% of that.
So at ICP3 volumes, price per decision is rarely the deciding factor. What decides is whether the decision is right, whether it arrives inside your turn budget at p95, and whether it gives you a probability you can threshold. If OpenAI prices Decisions far above Luna, or per call, rerun this table. The formula is:
monthly cost = calls x decisions_per_call x (input_tokens x input_rate + output_tokens x output_rate) / 1,000,000
Where image input matters for voice, and where it does not
Image input is the Decisions API's clearest confirmed advantage. Jev is text-only. For a pure phone line, this does not matter: a SIP call carries audio, and your decision model sees text from STT.
Images enter voice workflows in a few specific places:
- Screen-share or co-browse support. A web voice widget where the caller shares a screen. "Is the user on the billing settings page?" is a bounded question about an image.
- Document photo verification mid-call. The agent texts a link, the caller uploads a photo of an insurance card or ID, and the agent decides "legible / not legible / wrong document."
- Claims and field service. A photo of damage, a meter reading or a part label, classified into a fixed set before the agent continues.
- Receipt or order disputes. A photo of a receipt checked against "matches order / does not match / unreadable."
If none of these are on your roadmap, image input is not a reason to wait. If they are, you can run them on Jev today only by adding a vision model first to extract text, then deciding on the text. That adds a hop and its latency. A single Decisions call that takes the image directly could be simpler, once its contract and latency on image input are public. Image tokens cost more and may be slower, so do not reuse the 150 ms text figure for image decisions.
Probabilities, thresholds and abstention
This is the section most comparisons skip, and it matters most for voice.
A voice agent rarely wants just "the answer." It wants "the answer, and how sure." Escalation policy depends on it. If the agent is 97% sure the caller wants to cancel, it can route straight to the retention flow. If it is 55% sure, it should ask. If two routes are within a few points of each other, it should confirm. Without a probability, every decision is treated as equally certain, and you lose the ability to abstain.
The math of abstention
Geifman and El-Yaniv (2017), "Selective Classification for Deep Neural Networks" formalized this. Given a confidence score, you choose a threshold so the model answers only when confident, trading coverage (how often it answers) for risk (how often it is wrong when it does). On your own labeled data, you sweep the threshold and pick the point where error on answered cases meets your target.
For a voice agent, that becomes a three-band policy:
| Band | Rule (illustrative) | Agent action |
|---|---|---|
| Act | top probability at least 0.90 and margin over runner-up at least 0.20 | Route or run the tool |
| Confirm | top probability 0.60 to 0.90, or margin under 0.20 | Ask a short confirming question |
| Escalate or abstain | top probability under 0.60, or needs-human probability above your floor | Hand off or fall back |
Tune these on your own calls. Set tighter bands for high-cost routes (cancellations, payments) and looser bands for low-cost ones (store hours). Our routing and escalation guide walks through per-route thresholds, and escalation accuracy covers how to measure them.
Calibration: is 0.9 really 90%?
A probability is only useful if it is calibrated. Guo et al. (2017), "On Calibration of Modern Neural Networks" showed that modern deep networks are often overconfident, and that a simple post-hoc fix (temperature scaling) recovers much of the gap. The test is a reliability diagram: bucket predictions by stated confidence and check that the 0.9 bucket is right about 90% of the time. Expected Calibration Error (ECE) summarizes the gap as one number.
You do not need logits to check this. With returned probabilities and 500 to 1,000 labeled decisions, you can plot reliability yourself and, if needed, fit an isotonic mapping on top.
Option order is a real risk for any decision model
On Hacker News, a commenter (hbrn) reported that Jev's probabilities shift when the answer list is reordered, and that its confidence is derived from those probabilities. That critique is fair, and it is not unique to Jev. Research on LLMs shows the same effect:
- Pezeshkpour and Hruschka (2023) found performance gaps of roughly 13% to 75% across benchmarks when multiple-choice options were reordered. They traced it to positional bias when the model is torn between its top two or three choices, and got up to 8 points back with calibration methods.
- Zheng et al. (ICLR 2024) showed LLMs carry a prior toward specific option ID tokens and proposed PriDe, which estimates that prior by permuting options on a small sample and then removes it.
The practical rule for both products: fix your option order in code, never let it vary per request, and include an order-permutation test in your evaluation. If probabilities move a lot under permutation, average over two or three fixed orders for high-stakes decisions, or widen your Confirm band.
What to do if the Decisions API returns only a label
If OpenAI ships without probabilities, you have three workarounds, each with a cost:
1. Add an explicit "unsure" option. Cheap, but the model's use of "unsure" is itself uncalibrated.
2. Ask k times with permuted option order and treat agreement rate as confidence. This multiplies cost and, if run in series, latency.
3. Use it only where you do not need a threshold, such as post-call tagging reviewed in aggregate.
None of these are as good as a returned distribution for live escalation. If the Decisions API ships with probabilities, the gap closes, and you should rerun the calibration test on it.
Voice use-case map: which fits today
Not every decision has the same needs. The table maps the five most common voice decisions against what each product offers today.
| Use case | When it runs | Latency need | Needs probabilities | Images | Fit today |
|---|---|---|---|---|---|
| Routing and escalation | Every turn or first turns | Tight, blocking | Yes | No | Jev, documented and testable; Decisions after its p95 and probability support are published |
| Tool-call gating | Before each tool call | Tight, blocking | Yes, for confirm band | No | Jev; see tool calling with Jev |
| Disposition codes | Post-call | Loose | Helpful for review queues | Rarely | Either, once Decisions has pricing; Jev today |
| Compliance flags (disclosures, verification) | Live or post-call | Moderate | Yes (Noul-style probability) | No | Jev today; audit trail of probabilities helps |
| Live QA and call scoring | Post-call or near-live | Loose | Yes (score distributions) | Rarely | Jev today; see scoring every call |
| Photo or screen checks in a voice flow | Mid-call | Moderate | Helpful | Yes | Decisions API, once available to you; Jev needs a vision pre-step |
Two patterns stand out. First, the live, blocking decisions are where probabilities and tail latency matter most, and those are exactly the two things the Decisions API has not published. Second, post-call work is forgiving. If you already run on OpenAI and only need dispositions, waiting a week for Decisions docs costs you little.
For deeper treatment of the Jev side, see Jev for voice agents, Jev vs LLMs and call monitoring with Jev.
OpenAI's real advantages
A neutral comparison has to give OpenAI its due. Several advantages are real even before docs ship:
- Image input. Confirmed, and Jev does not have it.
- One vendor, one bill, one security review. If your voice stack already runs on OpenAI (GPT-Live, Realtime, or a GPT-6 model for the main LLM), adding Decisions avoids a new vendor contract, a new DPA and a new data-flow diagram. For a 50 to 500 person company with a lean security team, that can outweigh a price difference that is under 1% of call cost.
- Ecosystem and tooling. OpenAI's SDKs, dashboards, usage tracking, data residency options and enterprise tiers already exist. Decisions will likely plug into them.
- Model lineage. GPT-6 Luna is a general model with broad world knowledge. For decisions that need that knowledge, a Luna-based decider may do better on questions that are vague or depend on context outside the transcript. That is a hypothesis to test, not a fact.
Jev's advantages are equally concrete: it is available, documented, priced, returns probabilities for every option, is free on output, and has third-party latency telemetry. The LangChain study reported 100% agreement with human labels over 500 decisions, versus 99.8% for GPT-5.6 Terra, at $0.34 versus $28.17. That is one task, run by a partner, and it is not voice data. Use it as a reason to test, not as a result for your calls.
Reactions so far are mostly about what is missing. Firecrawl's analysis and Vercel's note both flag that OpenAI has not specified probability output. None of these are tests of the product.

Decision matrix: what to choose this week
Here is a reusable matrix. Find your row; the right column is the call we would make with today's public information.
| Your situation | Choose now | Why |
|---|---|---|
| Live routing or escalation on a phone line, need to ship this month | Jev, behind an adapter | Documented contract, probabilities, measured tail latency |
| Tool-call gating on payments or account changes | Jev, behind an adapter | Confirm band needs probabilities |
| Post-call dispositions only, stack already all-OpenAI | Wait for Decisions docs, then test both | Low latency need, procurement simplicity |
| Need image decisions (screen-share, document photo) | Plan for Decisions; prototype with a vision model plus Jev | Only Decisions takes images directly |
| Strict "single AI vendor" security policy | Decisions, once GA and tested | Avoids a new vendor review |
| Strict data-residency need | Check both vendors' current terms | Neither product's residency terms for this endpoint are covered here |
| You cannot get preview access | Jev now, Decisions as a planned challenger | You cannot test what you cannot call |
| High-stakes compliance flags with audit needs | Jev now; retest when Decisions publishes probabilities | Logged probabilities make audits defensible |
The principle behind every row: commit to an interface, not a vendor. The next section shows how.
A provider-agnostic decision adapter in Python
The cheapest insurance against an unfinished API is a thin adapter. Your agent code calls one function with a provider-neutral question. Each provider has a small class that translates it. When OpenAI publishes the Decisions contract, you fill in one class and run both through the same evaluation.
The Jev call below follows the TypeSafe quickstart: `TypeSafeClient`, `client.system_one(state=..., questions=...)`, `Choice(instructions=..., criteria=...)`, and `response.answers[name]` with `.choice`, `.confidence` and `.probabilities`. The OpenAI class is a placeholder on purpose. Its schema is unpublished, so we do not guess it.
# Simplified, illustrative adapter. Requires: pip install typesafe-sdk
# The client reads TYPESAFE_API_KEY from the environment.
import asyncio
import time
from dataclasses import dataclass, field
from typing import Optional, Protocol
from typesafe_sdk import Choice, TypeSafeClient
@dataclass(frozen=True)
class Question:
name: str
instructions: str
options: dict[str, str] # label -> description; keep order fixed in code
@dataclass
class Decision:
name: str
choice: str
probabilities: Optional[dict[str, float]] = None # None if provider lacks them
confidence: Optional[float] = None
provider: str = ""
latency_ms: float = 0.0
timed_out: bool = False
class DecisionProvider(Protocol):
name: str
def decide(self, state: str, questions: list[Question]) -> list[Decision]: ...
class JevProvider:
name = "jev"
def __init__(self) -> None:
self.client = TypeSafeClient()
def decide(self, state: str, questions: list[Question]) -> list[Decision]:
start = time.perf_counter()
response = self.client.system_one(
state=state,
questions={
q.name: Choice(instructions=q.instructions, criteria=q.options)
for q in questions
},
)
elapsed = (time.perf_counter() - start) * 1000
out = []
for q in questions:
a = response.answers[q.name]
out.append(Decision(
name=q.name,
choice=a.choice,
probabilities=dict(a.probabilities),
confidence=a.confidence,
provider=self.name,
latency_ms=elapsed,
))
return out
class OpenAIDecisionsProvider:
name = "openai-decisions"
def decide(self, state: str, questions: list[Question]) -> list[Decision]:
# TODO: fill in when OpenAI publishes the Decisions API contract
# (endpoint, request schema, response schema, SDK method).
# As of October 4, 2026 none of these are public. Do not guess them.
# Map the response to Decision objects. If no per-option probabilities
# are returned, leave probabilities=None so the policy below abstains
# from auto-acting.
raise NotImplementedError("OpenAI Decisions API contract not yet published")
async def decide_with_budget(
provider: DecisionProvider,
state: str,
questions: list[Question],
budget_ms: int,
fallback: dict[str, str],
) -> list[Decision]:
"""Run a blocking decision with a hard timeout so the caller never hears silence."""
try:
return await asyncio.wait_for(
asyncio.to_thread(provider.decide, state, questions),
timeout=budget_ms / 1000,
)
except (asyncio.TimeoutError, NotImplementedError):
return [
Decision(name=q.name, choice=fallback[q.name],
provider=provider.name, timed_out=True)
for q in questions
]
def policy(d: Decision, act_p: float = 0.90, min_margin: float = 0.20,
floor_p: float = 0.60) -> str:
"""Three-band policy. Illustrative thresholds; tune on your own labeled calls."""
if d.timed_out or d.probabilities is None:
return "confirm"
ranked = sorted(d.probabilities.values(), reverse=True)
top = ranked[0]
margin = top - (ranked[1] if len(ranked) > 1 else 0.0)
if top >= act_p and margin >= min_margin:
return "act"
if top < floor_p:
return "escalate"
return "confirm"
ROUTE = Question(
name="route",
instructions="Pick the flow that best serves the caller's main need on this turn.",
options={
"billing": "Charges, refunds, invoices or payment methods.",
"cancel": "Caller wants to end or downgrade a plan.",
"schedule": "Book, move or cancel an appointment.",
"support": "Something is broken: device, app or login.",
"other": "Anything else.",
},
)
async def on_user_turn(recent_turns: str, provider: DecisionProvider) -> str:
[route] = await decide_with_budget(
provider, recent_turns, [ROUTE], budget_ms=300,
fallback={"route": "other"},
)
action = policy(route)
# Log provider, choice, probabilities, latency_ms and action for every turn.
return f"{action}:{route.choice}"A few design notes that save pain later:
- Option order is fixed in code. The `options` dict preserves insertion order, which protects you from the order effects described above.
- The policy degrades safely. A provider that returns no probabilities, or a call that times out, lands in "confirm," never "act."
- Latency is measured at your boundary. `latency_ms` is what your agent experienced, which is the number you need, not the vendor's.
- The state is short. Pass recent turns and structured fields, not the full transcript, to stay clear of the irrelevant-state accuracy drop.
In a LiveKit agent server or a Pipecat pipeline, call `on_user_turn` from wherever you already handle the final user transcript, and start it as early as your end-of-turn signal allows.
How to test the OpenAI Decisions API vs Jev on your own calls
Vendor numbers will not tell you which product is right for your calls. A one-week, side-by-side test will. Our sibling post on benchmarking decision APIs for voice agents goes deeper on harness design; this is the short version.
1. Pick three to five decisions you actually make. For example: route, needs-human, tool gate for refunds, and post-call disposition. Write each as a fixed-order option list in the adapter format above.
2. Build a labeled set from real calls. Pull 300 to 1,000 decision points from recorded calls, stratified by route and including hard cases (mixed intents, angry callers, noisy transcripts). Have two people label each, and resolve disagreements. Our golden dataset guide covers sampling and labeling.
3. Freeze the state window. Decide exactly what each decision sees (for example, the last four turns plus account status) and use the same input for both providers.
4. Run Jev now, and Decisions as soon as you have access. Send identical inputs through the adapter. Record choice, probabilities (if any), and latency measured from your agent server's region.
5. Measure accuracy and calibration. Report accuracy per decision and per route. For any provider that returns probabilities, plot a reliability diagram and compute ECE. Check the Act band's error rate against your target.
6. Run the order test. Re-run 100 items with options in two other fixed orders. Count how often the chosen label changes and how far top probability moves.
7. Measure latency where it hurts. Report p50, p95 and p99 over at least 1,000 calls per provider, at your real concurrency, from your production region. Plug p95 and p99 into the turn budget and the at-least-one-slow-turn formula.
8. Price it on your traffic. Use your actual tokens per decision and decisions per call in the cost formula. Replace the Luna yardstick with real Decisions pricing the day it is published.
9. Decide with a written rule. For example: choose the provider with higher accuracy on high-stakes routes, provided its p95 stays under 300 ms and its Act-band error is under 2%. Write the rule before you look at results.
If you want this done independently, this is the kind of bake-off Evalgent runs: same calls, same inputs, both providers, with accuracy, calibration and tail latency reported against your own thresholds. The LLM-as-judge limits post explains why we do not score decision models with another LLM's opinion, and Jev as a judge covers when a decision model can itself be the scorer.
Frequently asked questions
Is the OpenAI Decisions API available to everyone?
Not as of October 4, 2026. OpenAI announced it at DevDay on September 29 as a limited preview for selected API customers, with broad release planned "in the coming days." There is no public docs page, schema or price yet. A third party, eesel, reported that a standard key received a 403 "not enabled" response.
How much does the OpenAI Decisions API cost?
OpenAI has not published Decisions API pricing. It is not on the API pricing page or the docs pricing table as of October 4, 2026. GPT-6 Luna lists at $0.10 per million input tokens and $0.50 per million output tokens, but that is a yardstick, not the Decisions price.
Does the Decisions API return probabilities or confidence scores?
OpenAI has not said. The DevDay text describes answers to questions with finite predefined answers, without mentioning probabilities. Firecrawl and Vercel both flagged this gap. Jev returns a probability for every option plus a confidence value, which is what escalation thresholds and abstention need.
Is 150 ms faster than Jev?
Not necessarily. The 150 ms figure is a DevDay slide comparing Decisions with OpenAI's own Luna API, with unstated conditions and no percentile. Jev's 0.21 s p50 and 0.34 s p95 are OpenRouter round-trip telemetry as of October 1. Measure both from your own servers before comparing.
Can Jev handle images?
No. Jev input is text only: a string, JSON object or array. For voice flows that involve a document photo or screen-share, you would run a vision model first to extract text, then decide with Jev. The Decisions API accepts images directly, which is its clearest confirmed advantage.
How consistent are Jev's answers?
Jev is highly consistent, though not perfectly fixed: the LangChain study found LLM judges' score variance 92 to 913 times higher than Jev's on its task. A Hacker News commenter showed probabilities shift when options are reordered, so fix option order in code and test permutations.
Which should I use for voice agent escalation today?
If you need to ship live escalation now, Jev is the one you can test against a documented contract, with per-option probabilities for a confirm band. Put it behind an adapter so you can trial the Decisions API when its docs, probability support and tail latency are published.
Will the cost difference matter at my call volume?
Probably not much. Under the assumptions in this post, Jev costs about $189 a month at 500,000 calls, and a Luna-priced equivalent would be $480 to $750. Both are under 1% of typical speech model cost per call. Accuracy, p95 latency and probabilities matter more.
The bottom line
Today, Jev is the decision model you can document, price, test and threshold for live voice decisions, while the OpenAI Decisions API is a credible challenger whose schema, probabilities, tail latency and price are still unpublished. Build behind a provider-agnostic adapter, test both on your own labeled calls when OpenAI opens access, and let accuracy, calibration and p95 latency inside your turn budget make the call.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more