Evalgent
Back to Blog
Voice AI Testing

Using OpenAI's Decisions API with GPT-Live in a Voice Agent: Client Delegation, Action Choices and Routing

Deepesh Jayal
19 min read
Using OpenAI's Decisions API with GPT-Live in a Voice Agent: Client Delegation, Action Choices and Routing
On this page

Last verified October 7, 2026, against OpenAI's Decisions guide, Connect voice to Decisions and Delegation and tools. The Decisions API is in public beta. OpenAI says GA is expected "in the coming weeks."

OpenAI's voice-to-Decisions guide shows the pattern with a browser ("reload this page") and a slide deck ("next slide"). Most teams reading this run phone agents. Their actions are order lookups, reschedules, SMS links and transfers. Those actions touch real records and can't always be undone. This post moves the pattern onto the phone. You get exact event names, an action catalog, server code, the safety checks the docs leave to you, a latency and cost model, and a test plan.

$0.10
Decisions input price per 1M tokens, no output charge (OpenAI Decisions guide)
500
Max tokens per commentary or thinking append (OpenAI delegation guide)
~10x
Decisions speed vs the Responses API, OpenAI's own claim with no published p95
$0.05
GPT-Live voice price per minute (OpenAI)

What the pattern is and why it works for voice

Two models share one call, and each does one job. GPT-Live listens, talks, handles turn-taking and fills the pause while work happens. The Decisions API picks one action from a fixed list you supply. Your server sits between them. It owns the state, runs the action and tells GPT-Live what happened.

OpenAI's guide sums it up: "GPT-Live continues speaking and listening while your app calls Decisions, runs the selected action, and returns the result." You never have to hide a slow backend behind silence. GPT-Live can say "let me pull that up" on its own while your code works.

Why not let a full LLM pick tools? There are three reasons, all from the docs.

1. Typed, bounded output. A `choice` answer is always one of your values, plus `probabilities` and a `confidence`. Your code can't receive a made-up tool name.

2. Speed. OpenAI says Decisions returns answers "about 10x faster than the Responses API." That is a vendor claim with no published p50 or p95. Measure it yourself (see the latency section below).

3. Price. It costs $0.10 per 1M input tokens with `gpt-6-luna`. There are "no cache-read, cache-write, or output-token charges."

The tradeoff is that Decisions doesn't generate arguments. It picks a label. Your app has to have the order number, date or phone number already. Or it picks `clarify` so GPT-Live asks for them. If you need a model that writes arguments or plans across steps, send that request to `reason` (covered below). If you are still comparing decision models, our OpenAI Decisions API vs Jev comparison covers that question. This post assumes you've picked Decisions and need to wire it in.

Responses delegationClient delegation + Decisions
Who picks the toolBackend Responses modelDecisions `choice`, then your code
OutputTool calls with generated argumentsOne label from your list, plus probabilities
Who keeps conversation contextLive manages backend stateYou collect transcripts and state
Switch mid-callNo. Needs a new sessionNo. Needs a new session
Best forOpen-ended multi-step tasksA known, finite action catalog

The "switch mid-call" row is a real gotcha. The delegation guide says changing modes needs a new Live session. Sending `delegation: null` to a running Responses session fails with `immutable_field_update`. Pick the mode before the call starts.

The event flow, step by step

Every name below comes from OpenAI's docs as of October 7, 2026.

1. Create the session in client mode. Use `{"model": "gpt-live-1", "delegation": {"type": "client"}}`. Your voice prompt should tell GPT-Live which actions the app supports, so it delegates at the right moments.

2. Collect transcripts. Listen for `session.input_transcript.delta` (caller) and `session.output_transcript.delta` (agent). Each has `delta` text plus `start_ms` and `end_ms`. The docs warn that "a transcript fragment is not a complete user turn, and transcripts may contain mistakes."

3. Receive the delegation. `session.delegation.created` carries `offset_ms` and a `delegation` object with `id` and `target: "client"`. It does not include what the caller asked for. Save `event.delegation.id` and send it back unchanged with every update about this task.

4. Build the Decisions input. Combine recent conversation, the latest request, current state and the actions available now. Input has to be a string or user-role messages. The API reference says "non-user roles, function calls, files, audio, and item references are not supported." So speaker labels go inside the text.

5. Call Decisions. `POST /v1/decisions` with `model: "gpt-6-luna"` and a `choice` question whose options include `noop`. In Python, that's `client.decisions.create(...)`. The SDK needs Python 3.26.0 or later.

6. Read and validate the answer. Find the answer by `name` in `decision.answers`. It's either `type: "choice"` (with `choice`, `confidence`, `probabilities`) or `type: "refusal"`. Skip the action if "the request was canceled or it no longer fits the current app state."

7. Execute. Run the action against your systems. Use an operation ID so a retry can't do it twice.

8. Return the result. Use `session.commentary.append` with `delegation_id` and `content` when GPT-Live should say it out loud ("it is trained to paraphrase the text"). Use `session.thinking.append` for facts it should know but not announce. Keep each append within 500 tokens. The acks are `session.commentary.appended` and `session.thinking.appended`. Match them to your `event_id` via `client_event_id`.

Swimlane sequence of one delegated voice action: caller speaks, GPT-Live emits transcript deltas and session.delegation.created, the server calls the Decisions API, validates and executes the action, and returns session.commentary.append while GPT-Live keeps talking

Two details here cause real bugs.

Cut the transcript at `offset_ms`. The delegation event has a timestamp but no text. If you build the prompt from "everything so far," you might include words the caller said after GPT-Live delegated. Usually that's harmless. Sometimes it's a correction, like "no, Thursday." Then your prompt mixes two requests. Use fragments with `start_ms <= offset_ms` for the decision. Treat later caller speech as a possible new revision of the request.

Acks don't mean the caller heard it. The docs say the ack "waits for estimated context injection, not for speech or playback to finish." The backend can finish while the spoken answer gets interrupted. Log both the action result and the output transcript, and test them separately.

From browser demos to phone calls: an action catalog

OpenAI's examples use `back`, `reload`, `noop` and slide controls. A phone agent needs a catalog where each action has a precondition your code can check. It also needs a reversibility flag and a channel for the result. Here is a copyable starting point for a support and scheduling line.

Action valueChoose when the caller…Precondition checked in codeReversible?Result channel
`order_status`asks where an order isCaller verified, order ID on fileRead-onlycommentary
`reschedule`wants a different time for an existing bookingVerified, booking exists, slot confirmed in stateYes, with carecommentary after success
`cancel_booking`wants to cancel a bookingVerified, booking exists, explicit confirmationNocommentary after success
`send_sms_link`wants a link texted (payment, tracking, form)Number on file or confirmed, consent recordedNo (message sent)commentary
`transfer_to_human`asks for a person, or the task is out of scopeBusiness hours or queue openNocommentary, then transfer
`clarify`gives an unclear or partial requestAlways availablen/ainstructions or commentary prompting one question
`reason`asks something that needs analysis or several stepsAlways availablen/acommentary from the reasoning model
`noop`is chatting, thanking, or no action fitsAlways availablen/anone, or thinking

Four design rules follow from the docs and from how classifiers fail.

Only offer actions that are open right now. OpenAI's browser example lists "Available actions: back, reload, noop" in the state. Do the same on calls. Remove `cancel_booking` from the `choices` array if the caller has no booking. A model can't pick an option that isn't in the request. It's cheaper to filter here than to reject the action afterward.

Give each option a distinct meaning. The Decisions guide says to "give choices distinct meanings." `noop`, `clarify` and `reason` sound alike, but they lead to different outcomes. `noop` means nothing to do. `clarify` means there's a request, but a required detail is missing. `reason` means a request your fixed actions can't cover. Write the `description` field so a human labeler could tell them apart.

Always include escape options. The docs' examples use `other` and `noop` for this. Without one, the model has to pick a real action even when none fits. We covered why forced answers inflate accuracy in our post on escape options and forced choice. The same logic applies to any decision model.

Split dependent questions. Put independent questions in one request. For example, a `choice` for the action plus a `predicate` for "the caller sounds like they want a person." Send dependent questions as separate requests. "Which slot did they confirm?" only makes sense after you know the action is `reschedule`.

For more phone routing patterns (escalation, disposition, compliance gating), see our voice routing use-case catalog and the guide to escalation accuracy.

Designing the Decisions input: small, current, ordered

The Decisions prompt is the only thing the model sees. OpenAI's voice guide uses a four-block layout: user conversation, last user request, current state, available actions. Keep that layout and add three fields for phone work: verification status, the pending confirmation (if any) and a task revision number.

User conversation (last 6 turns):
user: Hi, I need to move my appointment.
assistant: Sure. Which appointment is it?
user: The dental cleaning on Friday.

Last user request:
Can we do Thursday afternoon instead?

Current state:
Caller verified: yes. Booking B-2291: cleaning, Fri Oct 9 10:00.
Thursday Oct 8 open slots: 14:00, 15:30. No change made yet.
Pending confirmation: none. Task revision: 3.

Available actions: reschedule, cancel_booking, transfer_to_human, clarify, reason, noop.

Keep it small. Every token costs money and adds latency. Irrelevant context also hurts accuracy. Shi et al., in Large Language Models Can Be Easily Distracted by Irrelevant Context (ICML 2023), added irrelevant sentences to grade-school math problems. They found that "model performance is dramatically decreased when irrelevant information is included." Their tests were arithmetic, not routing. The production lesson still carries over: don't dump the full CRM record, every past call, or a raw tool payload into `input`. Include only the facts that change which action is right. The delegation guide gives the same advice for GPT-Live: "keep full HTML, DOM trees, large JSON payloads, and interaction logs in your application."

Keep it current. State comes from your systems, never from what the model said earlier. "Booking B-2291 exists" should be read from the database each time you build the prompt.

Keep the option order fixed. Zheng et al., in Large Language Models Are Not Robust Multiple Choice Selectors (ICLR 2024 Spotlight), tested 20 LLMs. They found the models "prefer to select specific option IDs," so answers shift when options move. Decisions isn't a plain LLM picking "A/B/C." It returns one of your values. Whether it has a position bias isn't published either way. So build `choices` in a stable order from one catalog, and include a permutation test in your test suite (see the test section below). The decision model benchmark protocol has a full order-sensitivity test.

Ballpark size. Six short turns, a state block and eight option descriptions come to a few hundred tokens. Log `decision.usage.input_tokens` from every response. Don't guess.

The server code (Python)

This is simplified but runnable in shape. It uses the methods shown in OpenAI's docs: `client.decisions.create`, `connection.session.commentary.append` and `connection.session.thinking.append`. It assumes you already have a Live connection or a sideband connection to the call. For SIP calls, that's the sideband WebSocket. Telephony setup is covered in our GPT-Live build guide. The business handlers are stubs.

# Simplified: GPT-Live client delegation + Decisions action routing.
# Verified against OpenAI docs (Oct 7, 2026). Handlers are stubs.
import asyncio, hashlib, json, time
from dataclasses import dataclass, field
from openai import AsyncOpenAI

client = AsyncOpenAI()
REASONING_MODEL = "your-reasoning-model"  # choose from OpenAI's reasoning models

CATALOG = {  # fixed order: never shuffle at runtime
    "order_status": "Look up where the caller's order is.",
    "reschedule": "Move an existing booking to a slot the caller confirmed.",
    "cancel_booking": "Cancel an existing booking the caller explicitly confirmed.",
    "confirm_pending": "Caller clearly said yes to the pending confirmation.",
    "decline_pending": "Caller said no or hesitated about the pending confirmation.",
    "send_sms_link": "Text the caller a link they asked for.",
    "transfer_to_human": "Caller asks for a person, or the task is out of scope.",
    "clarify": "A request exists but a required detail is missing or unclear.",
    "reason": "Request needs analysis or several steps none of the actions cover.",
    "noop": "No action fits, or the caller is just talking.",
}
MIN_CONFIDENCE = {"cancel_booking": 0.85, "confirm_pending": 0.85, "send_sms_link": 0.8}
IRREVERSIBLE = {"cancel_booking", "send_sms_link"}

@dataclass
class Call:
    caller_id: str
    verified: bool = False
    order_id: str | None = None
    booking: dict | None = None
    pending: str | None = None          # irreversible action awaiting a yes
    revision: int = 0
    turns: list = field(default_factory=list)   # [start_ms, role, text]
    done_ops: set = field(default_factory=set)

    def add(self, role, text, start_ms):
        if self.turns and self.turns[-1][1] == role:
            self.turns[-1][2] += text           # merge fragments of one turn
        else:
            self.turns.append([start_ms, role, text])

def available(c: Call) -> list[str]:
    if c.pending:
        acts = ["confirm_pending", "decline_pending"]
    else:
        acts = []
        if c.verified and c.order_id: acts.append("order_status")
        if c.verified and c.booking: acts += ["reschedule", "cancel_booking"]
        if c.verified: acts.append("send_sms_link")
    acts += ["transfer_to_human", "clarify", "reason", "noop"]
    return [a for a in CATALOG if a in acts]    # preserve catalog order

def build_input(c: Call, cutoff_ms: int) -> str:
    turns = [t for t in c.turns if t[0] <= cutoff_ms][-6:]
    convo = "\n".join(f"{role}: {text.strip()}" for _, role, text in turns)
    last = next((t[2] for t in reversed(turns) if t[1] == "user"), "")
    state = {"verified": c.verified, "order_id": c.order_id,
             "booking": c.booking, "pending_confirmation": c.pending,
             "task_revision": c.revision}
    return (f"User conversation:\n{convo}\n\nLast user request:\n{last}\n\n"
            f"Current state:\n{json.dumps(state)}\n\n"
            f"Available actions: {', '.join(available(c))}.")

async def decide(c: Call, cutoff_ms: int):
    acts = available(c)
    t0 = time.perf_counter()
    decision = await client.decisions.create(
        model="gpt-6-luna",
        input=build_input(c, cutoff_ms),
        questions=[{
            "type": "choice",
            "name": "action",
            "instructions": ("Choose the action the caller is requesting that is "
                             "available in the current state. Choose clarify if a "
                             "required detail is missing. Choose noop if no action fits."),
            "choices": [{"value": a, "description": CATALOG[a]} for a in acts],
        }],
        safety_identifier=hashlib.sha256(c.caller_id.encode()).hexdigest()[:64],
    )
    latency_ms = (time.perf_counter() - t0) * 1000
    answer = next(a for a in decision.answers if a.name == "action")
    print(json.dumps({"latency_ms": round(latency_ms),
                      "input_tokens": decision.usage.input_tokens,
                      "answer": answer.model_dump()}))
    return answer

async def handle_delegation(conn, c: Call, event):
    did, rev = event.delegation.id, c.revision
    answer = await decide(c, event.offset_ms)

    if answer.type == "refusal":
        await speak(conn, did, "I can't help with that here. Offer to connect a person.")
        return
    action = answer.choice
    if answer.confidence < MIN_CONFIDENCE.get(action, 0.6):
        action = "clarify"
    if c.revision != rev:                     # a newer request superseded this one
        return
    if action not in available(c):            # never trust the label alone
        action = "clarify"

    if action in IRREVERSIBLE and not c.pending:
        c.pending = action
        await speak(conn, did, confirm_prompt(c, action))
        return
    if action == "confirm_pending":
        action, c.pending = c.pending, None
    elif action == "decline_pending":
        c.pending = None
        await speak(conn, did, "Nothing was changed. Ask what they want to do instead.")
        return

    op_id = f"{c.caller_id}:{action}:{rev}"
    if op_id in c.done_ops:                   # retry or reconnect: don't do it twice
        return
    if action == "noop":
        return
    if action == "clarify":
        await speak(conn, did, "Ask one short question for the missing detail.")
        return
    if action == "reason":
        result = await reason(c, event.offset_ms)
    else:
        result = await HANDLERS[action](c)    # your systems; returns verified facts
    c.done_ops.add(op_id)
    if c.revision == rev:                     # drop results for outdated requests
        await speak(conn, did, result)

async def reason(c: Call, cutoff_ms: int) -> str:
    resp = await client.responses.create(
        model=REASONING_MODEL, reasoning={"effort": "low"},
        input=build_input(c, cutoff_ms) + "\n\nAnswer in two or three spoken sentences.")
    return resp.output_text

async def speak(conn, delegation_id: str, content: str):
    await conn.session.commentary.append(
        event_id=f"res_{delegation_id}_{int(time.time()*1000)}",
        delegation_id=delegation_id,
        content=content[:1600])   # rough guard for the 500-token append limit

def confirm_prompt(c: Call, action: str) -> str:
    return f"Before doing it, ask the caller to confirm: {action} for {c.booking}."

HANDLERS = {}  # "order_status": lookup_order, "reschedule": move_booking, ...

async def run(conn, c: Call):
    async for event in conn:
        if event.type == "session.input_transcript.delta":
            c.add("user", event.delta, event.start_ms)
        elif event.type == "session.output_transcript.delta":
            c.add("assistant", event.delta, event.start_ms)
        elif event.type == "session.delegation.created":
            c.revision += 1
            asyncio.create_task(handle_delegation(conn, c, event))

Some notes on the choices in this code:

  • `content[:1600]` is a rough cap based on about 4 characters per token, an assumption. Swap in a real tokenizer if your results run long. Better still, keep results short. The docs say "a concise tool result doesn't need an additional model call to rewrite it for speech."
  • `safety_identifier` is optional. The reference describes it as an opaque end-user ID of up to 128 characters, "never the authenticated user identity." A hash of the caller ID fits.
  • One `asyncio.create_task` per delegation lets a new request start while an older one is still running. The revision check is what keeps the stale one from speaking.
  • `clarify` uses commentary with an instruction-like sentence. GPT-Live paraphrases it into a question. If you need to redirect hard, for example after a guardrail block, the docs offer `session.instructions.append` with `delegation_id: null`. It "can interrupt the model's current speech."

Safety: the model picks, your code decides

The delegation guide is blunt: "your application owns permissions, confirmations, business records, and task state." Decisions picks a label. Five gates stand between that label and a changed record.

Decision tree for gating a Decisions API answer before execution: refusal check, confidence threshold, task revision freshness, availability and permission check, confirmation for irreversible actions, idempotent operation ID, then commentary or thinking append

1. Refusal. Any question can come back as `{"type": "refusal", "name": ...}`. Handle it on purpose. Don't let `answer.choice` throw an `AttributeError` mid-call. A good default is to offer a person.

2. Confidence threshold per action. The docs say to set thresholds "from labeled examples... based on the cost of false positives and false negatives." A wrong `order_status` costs one awkward sentence. A wrong `cancel_booking` costs a patient their slot. Use different thresholds for each. When the answer falls below the threshold, map it to `clarify`, not `noop`. The caller asked for something, so ask them about it.

3. Freshness. If a newer delegation has arrived, the old answer may be about "Friday" when the caller now means "Thursday." The docs: "use Thursday for subsequent work and ignore results from the outdated Friday request."

4. Availability and permission. Check the action against `available(state)` again just before running it. State can change between the prompt and the answer. For example, verification can expire or another channel can cancel the booking.

5. Confirmation for irreversible actions. Ask before you cancel, send or transfer. The code above does this with a two-step `pending` state. When a confirmation is pending, the only actions on offer are `confirm_pending` and `decline_pending`. That way the model can't jump to some other irreversible action at that moment.

Then make execution idempotent. OpenAI recommends tracking "each application action with its own operation ID and each changed request with a task revision." The guide also warns that "a lost response should not cause a second booking." The same goes for interruptions. "Interrupting the spoken conversation leaves backend work running." If the caller cuts in with "wait, stop," GPT-Live stops talking, but your reschedule call is still running. Handle cancellation in your backend, and say "canceled" only after the cancel actually succeeded.

If GPT-Live delegated an action, the fix is in your gates, not in GPT-Live's prompt. Our tool-calling gating guide goes deeper on pre-execution checks, and tool-call accuracy covers how to measure them.

Latency: a budget per hop, and how to measure it

There are two latencies, and they need separate numbers.

  • Acknowledgement latency. How long before the caller hears something. GPT-Live handles this on its own, because it keeps talking while you work. You don't need filler audio or "please hold" tricks.
  • Time to result. From the end of the caller's request to the moment GPT-Live starts speaking the verified result. This is the part your server owns.

OpenAI publishes no p50 or p95 for Decisions, only the "about 10x faster than the Responses API" claim. The table below is a worked example with assumed numbers. Replace each one with your own p95.

HopAssumed p95 (illustrative)How to measure
Caller stops speaking → `session.delegation.created` at your server300 msLast `input_transcript.delta` `end_ms` vs your receive time, aligned with `offset_ms`
Build prompt from state10 msTimer around `build_input` (plus any DB reads)
`client.decisions.create` round trip250 ms`perf_counter` around the call, logged per request
Gates and validation5 msTimer
Business action (order lookup, booking API)400 msTimer per handler, tagged by action
`session.commentary.append` → `session.commentary.appended`100 msMatch `client_event_id` to your `event_id`
Ack → first `output_transcript.delta` of the result300 msFirst output delta after the ack
Time to result (sum of p95s, worst-case style)≈1,365 msEnd-to-end per delegation

Adding p95s together overstates the true p95 of the total, so read the sum as a rough upper bound. The useful lesson is the shape. In this example, Decisions is a minority of the time to result, and the business action is the biggest hop. Check whether that's true for your stack. If the Decisions hop is small, speeding it up won't help much. Speed up the slowest backend action instead.

Per-hop latency waterfall for one delegated phone action with illustrative p95 values, showing GPT-Live acknowledgement speech overlapping the Decisions call and business action until the spoken result begins

There are two ways to shorten time to result.

Start early. The delegation guide describes processing transcript fragments "to start work before a delegation event arrives." That includes speculative lookups. A read-only `order_status` lookup is safe to prefetch once an order number shows up in the transcript. An irreversible action never is. If you prefetch, the docs ask you to "discard outdated results" and "avoid duplicate actions."

Return progress quietly. For a slow action, send `session.thinking.append` with something like "Checking Thursday availability. No appointment has been booked." GPT-Live can then answer "is it done yet?" truthfully without announcing every step.

Cost per call: the decision layer is a rounding error

Pricing is $0.10 per 1M input tokens with `gpt-6-luna`, with no output charge. Regional processing premiums and long-context multipliers apply, so check them if you pin data to the EU.

Labeled assumptions for a typical support call:

  • 600 input tokens per decision (6 turns, state, 8 option descriptions). This is an assumption. Log `usage.input_tokens` to get your real number.
  • 4 delegations per call.
  • 4-minute call on GPT-Live at $0.05 per minute.

Arithmetic:

  • Decisions per call: 600 × 4 = 2,400 tokens × $0.10 / 1,000,000 = $0.00024
  • GPT-Live per call: 4 × $0.05 = $0.20
  • At 100,000 calls a month: Decisions ≈ $24, GPT-Live ≈ $20,000

Even if the prompt grows to 1,500 tokens and you make 6 decisions per call, that's 9,000 tokens, or $0.0009 per call ($90 a month at 100,000 calls). The real cost lever is the `reason` route. It calls the Responses API with a reasoning model, billed at that model's input and output rates. Track the share of delegations that go to `reason`. A rising share is both a cost signal and a sign that your action catalog is missing something callers want.

The `reason` route for complex requests

OpenAI's slide example adds a `reason` choice: "Use a reasoning model for analysis, planning, or other requests." "Go to the next slide" maps to `next_slide`. "Compare these two plans and recommend one" maps to `reason`. On a phone line, `reason` catches requests like "which of my two orders will get here first, and can you combine them?" or "explain why my bill went up." No single fixed action answers those.

Three rules keep it under control:

1. Keep the same delegation ID. The docs say to keep the session in client mode for both routes. Return the reasoning result "with the same delegation ID."

2. Pass only what's relevant. That means the original request and the specific records needed, not the whole account.

3. Treat reasoning output as advice, not action. If the reasoning model concludes "cancel order A," don't execute it directly. Put it back through the gates: a fresh decision, a confirmation and the operation ID check. The `reason` route is where unvalidated actions tend to slip in.

`reason` is also a test target. If an easy request like "where's my order?" lands on `reason`, that's a routing error. It costs money and adds seconds even when the final answer is right.

Putting it on a phone line

GPT-Live accepts calls over SIP directly. Your carrier routes the call to OpenAI, your server accepts it, and you attach a sideband WebSocket. Your server runs the delegation handler shown above on that sideband, so it never touches audio. See OpenAI's Telephony and SIP guide for the flow. Our SIP vs media-stream bridge guide compares native SIP, a Twilio or Telnyx bridge, and LiveKit or Pipecat plugins.

Two phone-specific points matter for routing:

  • Phone transcripts are noisier. 8 kHz audio, accents and background noise make "Thursday" vs "Tuesday" and order numbers less reliable. When an entity matters, have `clarify` read it back. Get the caller's confirmation before an action that depends on it.
  • `transfer_to_human` ends the delegation loop. Send the commentary ("connecting you now") and wait for the ack. Then do the transfer using your SIP transfer path. Test that the caller hears the handoff line before the audio drops.

How to test a Decisions-plus-GPT-Live action router

Testing has two layers. The decision layer is text in, label out, and runs fast. The call layer is audio in, action plus speech out, and runs slower. Do both. Our GPT-Live testing guide covers the audio harness. The steps below cover the routing layer.

1. Build a labeled set from real call snippets. Cut each sample at the delegation point, using the same `build_input` your server uses. Label the correct action. Aim for at least 200 samples per action. For 95% expected accuracy with a ±3-point 95% interval, n = 1.96² × 0.95 × 0.05 / 0.03² ≈ 203. Oversample `noop`, `clarify` and `reason`, since they're the easiest to confuse.

2. Measure accuracy per action, not overall. One overall number hides a bad `cancel_booking` behind a good `order_status`. Report a confusion matrix.

3. Measure noop precision and recall. Noop recall = correct noops ÷ all true noops. Low recall means the agent acts when it shouldn't. Noop precision = correct noops ÷ all predicted noops. Low precision means the agent ignores real requests. Track `clarify` the same way.

4. Set per-action thresholds from the data. Sweep `confidence` and pick, for each action, the lowest threshold that meets your error budget. Re-check after any change to the prompt or catalog.

5. Run stale-state tests. Script calls where state changes between delegation and execution. Examples: the booking is canceled on the web, verification expires, the caller corrects "Friday" to "Thursday." Pass means zero actions run on outdated state and zero outdated results get spoken.

6. Run interruption tests. Barge in during the business action ("wait, stop"). Then check the record and the transcript separately. The pass condition is no "done" or "canceled" claim unless the backend confirms it.

7. Run an option-order test. Shuffle `choices` order on a copy of the test set and compare labels. Any flip rate above zero on clear cases needs a look. Keep the production order fixed either way.

8. Measure latency at p95 by hop. Use the table above as your log schema. Gate releases on time-to-result p95, not the mean.

9. Measure pass^k on full calls. τ-bench (Yao et al., 2024) introduced pass^k: the chance that all k trials of a task succeed. The authors found strong function-calling agents were "quite inconsistent (pass^8 <25% in retail)." For a call scenario with c successes in n trials, estimate pass^k = C(c,k) / C(n,k). Even 95% per-run success gives about 0.95^8 ≈ 66% pass^8, if runs are independent.

Suggested starting thresholds for a support line. These are assumptions to tune, not standards.

MetricStarting pass bar
Accuracy, read-only actions≥ 95%
Accuracy, irreversible actions≥ 98%, plus confirmation in 100% of executions
Noop recall≥ 95%
Actions executed on stale state0
False "done" claims after interruption0
Time-to-result p95Your product target, measured per action
pass^8 on top 20 scenariosTrack the trend; investigate any drop

For the full benchmark protocol (calibration, Wilson intervals, repeatability), use our decision model benchmark guide. Without an eval team, this is where an independent audit pays off. Evalgent can run the scripted stale-state, interruption and pass^k suites before launch, and again after each prompt or catalog change.

Frequently asked questions

What is GPT-Live client delegation?

It is a GPT-Live session mode set with `delegation: {type: "client"}`. GPT-Live handles the spoken conversation. When it needs work done, it emits `session.delegation.created` with a delegation ID. Your application runs the work with any model or service and returns results with `session.commentary.append` or `session.thinking.append`, using the same ID.

Does the delegation event include what the caller asked for?

No. `session.delegation.created` contains `offset_ms` and delegation metadata (`id`, `target`), but no task text. You rebuild the request from `session.input_transcript.delta` and `session.output_transcript.delta` events plus your app state. Use `offset_ms` to cut the transcript at the moment GPT-Live delegated.

When should I use commentary vs thinking appends?

Use `session.commentary.append` for results GPT-Live should say out loud, such as "your order shipped today." It paraphrases the text. Use `session.thinking.append` for progress or facts it should know without announcing them. Both take plain-string content up to 500 tokens and a `delegation_id`.

How much does the Decisions API cost per call?

$0.10 per 1M input tokens with `gpt-6-luna`, with no output, cache-read or cache-write charges. With assumed 600-token prompts and four decisions per call, that's $0.00024 per call. That is far below a 4-minute GPT-Live call at $0.05 per minute. Regional premiums and long-context multipliers apply.

How fast is the Decisions API?

OpenAI says "about 10x faster than the Responses API." It publishes no p50, p95 or rate limits. Measure it with a timer around `client.decisions.create` on your own prompts and region, and track p95 per hop. In many phone stacks the business action, not the decision, is the slowest hop.

Why include a noop option?

Without an escape option, a choice model must pick a real action even when the caller is just chatting or saying thanks. OpenAI's voice example tells the model to "Choose noop if no action fits." Add `clarify` for requests that are missing details, and measure noop precision and recall separately.

Can the Decisions API take audio directly?

No. Input must be a string or user-role messages with `input_text` and inline `input_image` parts. The API reference says audio, files, tools and non-user roles are not supported. In a voice agent you pass GPT-Live's transcripts as text, with speaker labels inside the string.

What happens if Decisions refuses a question?

Any question can return `{"type": "refusal", "name": ...}` instead of an answer. Check `answer.type` before reading `answer.choice`. A safe default on a phone line is to send commentary that offers a human transfer, and to log the refusal for review.

The bottom line

With client delegation, the Decisions API and GPT-Live split one call cleanly. GPT-Live keeps the conversation going, a bounded choice picks the action, and your code alone decides whether that action runs. Ship it with an action list filtered by current state, escape options, per-action thresholds, freshness and confirmation gates. Then test routing accuracy, noop behavior, stale state, interruptions and pass^k before real callers find the gaps.

Related Articles