Evalgent
Back to Blog
Voice AI Testing

OpenAI Realtime API SIP vs a Media-Stream Bridge: Putting OpenAI Realtime, GPT-Live, or Gemini Live on a Phone Line

Deepesh Jayal
24 min read
OpenAI Realtime API SIP vs a Media-Stream Bridge: Putting OpenAI Realtime, GPT-Live, or Gemini Live on a Phone Line
On this page

You built a speech-to-speech agent that sounds great in the browser. Now it has to answer a phone number. That is where most realtime projects break. The phone leg adds an 8 kHz codec, a carrier-side playback buffer you cannot see, and call-control actions the model API never knew about.

This guide covers the wiring. If you are still choosing a model, start with our guide to evaluating realtime voice APIs. For GPT-Live specifics such as delegation and prompts, see building with GPT-Live. For the test strategy behind speech-to-speech models in general, read testing speech-to-speech voice agents.

Everything below was checked against the OpenAI, Google, Twilio, LiveKit, and Pipecat docs on October 4, 2026. Model names in use today are `gpt-realtime-2.1` and `gpt-live-1` on OpenAI and `gemini-3.8-live` on Google.

60 min
maximum OpenAI Realtime session length (OpenAI docs)
15 min
Gemini Live audio-only session cap without compression (Google docs)
~10 min
Gemini Live connection lifetime before a reset (Google docs)
8 bytes
of G.711 mu-law audio per millisecond at 8 kHz (Twilio stream format)

The three architectures, in one paragraph each

(a) Native SIP. Your SIP trunk sends the call straight to OpenAI. OpenAI fires a webhook, and you accept or reject it. Then you attach a "sideband" WebSocket to watch events, run tools, and send commands. Your server never touches audio. This works for the Realtime API and for GPT-Live. Google's Gemini API docs list no SIP endpoint as of this writing.

(b) Media-stream bridge. Twilio or Telnyx streams call audio to your WebSocket server as base64 JSON frames. Your server opens a second WebSocket to the model and relays audio both ways. You own every hard part: codec conversion, the carrier's playback buffer, barge-in, transfers, and hangups.

(c) Framework plugin. The call enters LiveKit SIP or a Pipecat transport. A plugin like LiveKit's `openai.realtime.RealtimeModel` or Pipecat's `GeminiLiveLLMService` talks to the model. The framework owns playout tracking and interruption logic. You get its fixes and you inherit its bugs.

Three call paths compared: native SIP trunk to OpenAI with a sideband WebSocket, a Twilio media-stream bridge through your server, and a LiveKit or Pipecat framework plugin, showing hops and who owns the playback buffer

Path A: OpenAI's native SIP endpoint, step by step

The Realtime SIP guide describes the flow. It is short, but each step hides a production decision.

1. Create a project webhook for incoming calls in the OpenAI platform settings.

2. Point your SIP trunk at `sip:$PROJECT_ID@sip.api.openai.com;transport=tls`.

3. Receive the `realtime.call.incoming` webhook. It carries a `call_id` and the SIP headers.

4. `POST /v1/realtime/calls/{call_id}/accept` with your session config (model, voice, tools, instructions). Or reject with a SIP status such as 486. The default is 603 Decline.

5. Open `wss://api.openai.com/v1/realtime?call_id={call_id}` to watch events and send commands.

6. Use `/refer` to transfer and `/hangup` to end the call.

Here is a minimal handler, simplified from OpenAI's Python example. It adds two things the example leaves out: webhook deduplication and a dedicated sideband task per call.

# Simplified, illustrative. Based on OpenAI's realtime-sip Python example (Oct 2026).
import asyncio, json, os
import httpx, websockets
from fastapi import FastAPI, Request, Response
from openai import OpenAI, InvalidWebhookSignatureError

app = FastAPI()
client = OpenAI(webhook_secret=os.environ["OPENAI_WEBHOOK_SECRET"])
AUTH = {"Authorization": f"Bearer {os.environ['OPENAI_API_KEY']}"}
seen_webhooks: set[str] = set()   # use Redis with a TTL in production

ACCEPT = {
    "type": "realtime",
    "model": "gpt-realtime-2.1",
    "instructions": "You answer inbound calls for Acme Dental. Be brief.",
    "audio": {"output": {"voice": "marin"}},
    "tools": [],  # function schemas go here
}

@app.post("/openai/webhook")
async def webhook(request: Request):
    body = await request.body()
    try:
        event = client.webhooks.unwrap(body, request.headers)
    except InvalidWebhookSignatureError:
        return Response(status_code=400)
    wid = request.headers.get("webhook-id")
    if wid in seen_webhooks:            # retries happen; never accept twice
        return Response(status_code=200)
    seen_webhooks.add(wid)
    if event.type == "realtime.call.incoming":
        call_id = event.data.call_id
        async with httpx.AsyncClient() as h:
            r = await h.post(f"https://api.openai.com/v1/realtime/calls/{call_id}/accept",
                             headers=AUTH, json=ACCEPT)
        if r.status_code == 200:
            asyncio.create_task(sideband(call_id))
    return Response(status_code=200)

async def sideband(call_id: str):
    url = f"wss://api.openai.com/v1/realtime?call_id={call_id}"
    async with websockets.connect(url, additional_headers=AUTH) as ws:
        await ws.send(json.dumps({"type": "response.create"}))  # speak first
        async for raw in ws:
            evt = json.loads(raw)
            if evt["type"] == "input_audio_buffer.dtmf_event_received":
                print("keypad", evt["event"])       # SIP-only event
            # route function calls, log usage from response.done, etc.

Four details matter on a real trunk:

  • Network allowlists. `sip.api.openai.com` and `sip-eu.api.openai.com` are GeoIP-routed. Signaling uses TLS on port 5061. SRTP media comes from separate CIDRs listed in the docs. A firewall that allows signaling but not those media ranges produces a classic one-way-audio call.
  • Accept fast. OpenAI returns 200 once the SIP leg is ringing and the session is being set up. Your webhook handler should not run slow lookups before accepting. Do caller lookups after accept, then inject context with `session.update`.
  • Barge-in is handled for you. OpenAI owns the media, so it knows what was played. You do not compute truncation offsets on this path. That alone removes the most common bridge bug, covered below.
  • DTMF is an event, not audio. Key presses arrive as `input_audio_buffer.dtmf_event_received` on the sideband.

GPT-Live on SIP is a different API

GPT-Live uses the same idea with different names. The GPT-Live telephony guide uses a `live.transport.incoming` webhook with `data.session_id`. You accept with `POST /v1/live/sessions/{session_id}/accept` and attach at `/v1/live/sessions/{session_id}/attach`. DTMF arrives as `transport.dtmf.received`.

Three GPT-Live rules are easy to miss:

  • The first decision wins. A later competing accept or reject returns `decision_already_made`. The same pending call can also fire a Realtime webhook, so assign exactly one handler.
  • Outbound SIP exists, but only on the Live API. It must be enabled for your organization. Ringing is capped at 3 minutes and a connected call at 2 hours.
  • Never auto-retry an outbound create. The docs say `X-Client-Request-Id` does not deduplicate. A retry after a timeout can place a second call to the same person. For a collections or reminders agent, that is a compliance incident, not a bug.

Attach the sideband right away on outbound calls. It only replays the previous 3 seconds of events, so a late attach can miss `transport.ringing` or `transport.answered`.

Path B: the media-stream bridge, message by message

The bridge is the path most GitHub repos use, including Twilio's own OpenAI Realtime sample. Our Pipecat Twilio and Telnyx guide covers the carrier side in depth. Here is what changes when the far end is a speech-to-speech model.

Twilio's Media Streams protocol is fixed: `audio/x-mulaw`, 8000 Hz, mono, base64 in JSON. Twilio sends `connected`, `start`, `media`, `dtmf`, `mark`, and `stop`. You send `media`, `mark`, and `clear`. Twilio buffers your outbound media and plays it in order. You cannot see that buffer except through `mark` echoes.

Codec passthrough versus resampling

Each model accepts different audio. This table decides how much DSP your bridge does.

ModelInput formats on a WebSocketOutputBridge work for Twilio
OpenAI Realtime (`gpt-realtime-2.1`)`audio/pcm` at 24 kHz or `audio/pcmu`Same choicesNone. Set both sides to `audio/pcmu` and forward bytes
GPT-Live (`gpt-live-1`)Raw G.711 mu-law or A-law at 8 kHzSameNone if codec, rate, and channels match
Gemini Live (`gemini-3.8-live`)16-bit PCM; native 16 kHz, any rate accepted via `audio/pcm;rate=N`16-bit PCM, always 24 kHzDecode mu-law in; downsample 24k to 8k and encode out

Two points are not obvious.

First, upsampling phone audio adds no information. A G.711 call is band-limited to roughly 300 to 3,400 Hz before it reaches you. Converting it to 24 kHz PCM16 does not restore the missing band. It only adds CPU time and a filter delay. On OpenAI, `audio/pcmu` passthrough is the right default. On Gemini, the capabilities guide says the API resamples any input rate. So you can decode mu-law to 8 kHz PCM16 and send it with `audio/pcm;rate=8000`. Let Google resample instead of running your own upsampler.

Second, the output path needs a real low-pass filter. Gemini always returns 24 kHz. Going to 8 kHz is a clean factor of 3, but dropping two of every three samples aliases everything above 4 kHz back into the band. Sibilants turn into a metallic hiss. Use a polyphase resampler such as `scipy.signal.resample_poly(x, 1, 3)` and then mu-law-encode. Also note that Python's `audioop` module was removed in Python 3.13. Many bridges that used `audioop.lin2ulaw` break on upgrade. Pin `audioop-lts` or use a NumPy lookup table.

The bandwidth math

Bandwidth matters once you run dozens of concurrent calls on one bridge host.

  • Twilio leg: 8,000 bytes per second of mu-law, about 10,700 bytes per second after base64. That is roughly 86 kbps per direction.
  • Gemini output leg: 24,000 samples x 2 bytes = 48,000 bytes per second, or 64,000 after base64. That is about 512 kbps while the agent speaks.
  • At 100 concurrent calls with the agent talking, Gemini output alone is 100 x 512 kbps, or about 51 Mbps into the bridge.

The OpenAI `audio/pcmu` path stays at telephone bitrates end to end. That is about six times less bridge traffic than a 24 kHz PCM16 output path.

Barge-in: the event mapping that breaks most bridges

Barge-in on a bridge is a three-party problem. The model generates audio faster than real time. Your bridge forwards it. Twilio buffers it and plays it at real time. When the caller interrupts, three clocks disagree about what the caller heard.

OpenAI Realtime: speech_started, clear, truncate

The correct sequence on the Realtime API is:

1. The server sends `input_audio_buffer.speech_started`. With `interrupt_response: true` (the default), the server also cancels the in-flight response.

2. Your bridge sends Twilio `{"event": "clear", "streamSid": ...}` to dump buffered audio.

3. Your bridge sends `conversation.item.truncate` with the assistant `item_id`, `content_index: 0`, and `audio_end_ms`.

`audio_end_ms` must equal the audio the caller actually heard. The client events reference says truncation also deletes the server-side transcript past that point. That keeps text the caller never heard out of the context. If `audio_end_ms` is larger than the item's audio, the server returns an error and nothing is truncated.

Twilio's sample computes `audio_end_ms` as the latest inbound media timestamp minus the timestamp when the first outbound delta was forwarded. That is wall-clock time since sending started, not played audio. A January 2026 OpenAI community thread describes the result. Truncation lands at the wrong point and the conversation state drifts.

Wall-clock estimates fail in three ways:

  • Late start. Twilio's playout starts after network and jitter delay, so elapsed time overcounts.
  • Generation gaps. If audio deltas pause mid-response, the clock keeps running while nothing plays. The estimate can exceed the audio sent, and the truncate errors out.
  • Clear echoes. Twilio's docs say a `clear` makes Twilio "send back mark messages matching any remaining mark messages." Code that treats every mark echo as "played" counts cleared audio as heard.
Barge-in timeline on a Twilio bridge: model audio sent faster than real time, Twilio playout lagging behind, the caller interrupting, and a wall-clock truncation estimate overshooting what was heard while a mark-based ledger lands on the true point

The fix is a played-audio ledger. Name each mark with the cumulative milliseconds of audio sent for that item. When Twilio echoes a mark, everything up to that offset has played. Freeze the ledger before you send `clear`, so cleared-audio echoes are ignored.

# Simplified, illustrative: played-audio ledger for a Twilio <-> OpenAI Realtime bridge.
import base64, json, time

MS_PER_BYTE = 1 / 8          # G.711 mu-law at 8 kHz: 8 bytes per ms

class PlaybackLedger:
    def __init__(self):
        self.reset(None)

    def reset(self, item_id):
        self.item_id, self.sent_ms, self.acked_ms = item_id, 0, 0
        self.acked_at, self.frozen = None, False

    def on_audio_delta(self, item_id: str, b64: str) -> str:
        if item_id != self.item_id:
            self.reset(item_id)
        self.sent_ms += int(len(base64.b64decode(b64)) * MS_PER_BYTE)
        return f"{item_id}|{self.sent_ms}"   # use as the mark name

    def on_mark(self, name: str):
        item_id, ms = name.rsplit("|", 1)
        if self.frozen or item_id != self.item_id:
            return                           # cleared-audio echo: not played
        self.acked_ms, self.acked_at = max(self.acked_ms, int(ms)), time.monotonic()

    def heard_ms(self) -> int:
        if self.acked_at is None:
            return 0                         # first chunk still playing
        since = int((time.monotonic() - self.acked_at) * 1000)
        return min(self.sent_ms, self.acked_ms + since)

async def on_speech_started(openai_ws, twilio_ws, stream_sid, ledger):
    heard = ledger.heard_ms()
    ledger.frozen = True                     # freeze BEFORE clear
    await twilio_ws.send_text(json.dumps({"event": "clear", "streamSid": stream_sid}))
    if ledger.item_id and ledger.sent_ms:
        await openai_ws.send(json.dumps({
            "type": "conversation.item.truncate",
            "item_id": ledger.item_id,
            "content_index": 0,
            "audio_end_ms": heard,
        }))

The session config for passthrough uses `audio/pcmu` both ways:

{
  "type": "session.update",
  "session": {
    "type": "realtime",
    "model": "gpt-realtime-2.1",
    "output_modalities": ["audio"],
    "audio": {
      "input":  {"format": {"type": "audio/pcmu"},
                 "turn_detection": {"type": "semantic_vad", "interrupt_response": true}},
      "output": {"format": {"type": "audio/pcmu"}, "voice": "marin"}
    }
  }
}

The ledger's error is bounded by one delta's duration, because the first mark of a response has not echoed yet. Smaller outbound chunks tighten it. You can split large deltas into 100 ms pieces before forwarding, which is 800 bytes of mu-law each.

Gemini Live: interrupted, with no truncate

Gemini works differently. The capabilities guide says that when VAD detects an interruption, "the ongoing generation is canceled and discarded. Only the information already sent to the client is retained in the session history." The server then sends `serverContent.interrupted` and the IDs of any canceled function calls.

On a phone bridge, "sent to the client" means sent to your bridge. Audio sitting in Twilio's buffer counts as delivered. So Gemini's history includes words the caller never heard, and there is no truncate event to fix it. That is the same drift the Realtime truncate exists to prevent.

You have two levers:

  • Keep the carrier buffer shallow. Pace outbound audio to Twilio at real time plus a small lead, rather than dumping every chunk on arrival. Hold the rest in your own queue, which you can drop instantly on `interrupted`. This does not change Gemini's view, but it shrinks how much the caller misses after `clear`.
  • Tell the model what was heard. `gemini-3.8-live` supports `send_client_content` throughout the session with explicit roles. Without `turn_complete`, the server waits instead of responding. After an interruption, you can send a short note that the caller heard only up to a given phrase. Use output transcription and your ledger to find that phrase. Treat this as a pattern to validate on your own calls, not a documented feature.

The echo self-interrupt

A May 2026 Google developer forum report describes a telephony-specific failure. Some carriers reflect the model's outbound audio back into the inbound stream. Gemini's VAD hears its own words, fires `interrupted`, and the transcript attributes the model's words to the caller. The reporter's workaround, `prefixPaddingMs: 200`, filtered short echoes but not longer phrases.

This is not Gemini-specific in principle. Any full-duplex model with server-side VAD can be tricked by line echo or speakerphone bleed. Detect it by comparing input transcripts to the agent's recent output transcript. Then fix it at the media layer: echo cancellation on the inbound leg, or a softer start-of-speech sensitivity. Our guide to noise and echo cancellation on self-hosted LiveKit covers the media-layer options.

GPT-Live on a bridge

GPT-Live does not support truncation, and it has no per-response done event. The migration guide says to drive the speaking indicator from your audio player. On a bridge, that player is Twilio's buffer, which you only observe through marks. Pacing matters even more here. When `session.input_transcript.delta` arrives while your ledger shows queued audio, send `clear` and drop your local queue.

The mapping table

Event or actionOpenAI Realtime (bridge)GPT-Live (bridge)Gemini Live (bridge)Native SIP (OpenAI)
Caller starts talking`input_audio_buffer.speech_started`Input transcript deltas during playback`serverContent.interrupted`Handled server-side
Stop generationAutomatic with `interrupt_response`Model decidesAutomatic, generation discardedAutomatic
Flush carrier audioTwilio `clear`Twilio `clear`Twilio `clear`Not needed
Fix context`conversation.item.truncate` with heard msNot supportedNot supported; optional client noteServer knows playout
Pending tool callsYour code decidesBackend task trackingCanceled IDs sent by serverYour sideband decides

Scoring this behavior is its own discipline. Our guide to evaluating interruption detection defines the metrics.

Session limits and reconnection

Phone calls are long-lived connections. Each provider caps them differently.

LimitOpenAI RealtimeGPT-LiveGemini Live
Session cap60 minutes2 hours connected (outbound SIP)15 min audio-only without compression
Connection resetNot documentedNot documentedAbout 10 minutes, with a `GoAway` warning
ExtensionNew session plus summarized contextNew sessionContext window compression plus session resumption
Resume tokenNot applicableNot applicableValid 2 hours after the last session ends

Sources: the Realtime conversations guide, the GPT-Live telephony guide, and the Gemini session management guide.

The Gemini numbers change your architecture. Any call longer than about 10 minutes will cross at least one connection reset. A support call that runs 14 minutes is normal. Without resumption, the session dies mid-sentence and the caller hears silence.

The handling pattern looks like this:

# Simplified, illustrative: Gemini Live reconnect loop for a phone bridge.
import asyncio
from google import genai
from google.genai import types

client = genai.Client()

def live_config(handle):
    return types.LiveConnectConfig(
        response_modalities=["AUDIO"],
        session_resumption=types.SessionResumptionConfig(handle=handle),
        context_window_compression=types.ContextWindowCompressionConfig(
            sliding_window=types.SlidingWindow()),
        input_audio_transcription=types.AudioTranscriptionConfig(),
        output_audio_transcription=types.AudioTranscriptionConfig(),
    )

async def run_call(call):
    handle = None
    while call.active:
        async with client.aio.live.connect(model="gemini-3.8-live",
                                           config=live_config(handle)) as session:
            sender = asyncio.create_task(call.pump_caller_audio(session))  # replays buffered audio first
            try:
                async for msg in session.receive():
                    u = msg.session_resumption_update
                    if u and u.resumable and u.new_handle:
                        handle = u.new_handle            # always keep the latest
                    if msg.go_away is not None:
                        call.start_buffering_caller_audio()
                        break                            # reconnect with handle
                    if msg.server_content and msg.server_content.interrupted:
                        await call.flush_playback()      # Twilio clear + local queue
                    if msg.data:
                        await call.play_24k_pcm(msg.data)
            finally:
                sender.cancel()

Three rules make this reliable:

  • Buffer caller audio during the handshake. Keep a short ring buffer while you reconnect and replay it on the new connection. Otherwise the first words after the reset are lost.
  • Reconnect between turns when you can. `GoAway` carries `timeLeft`. If the model is mid-utterance, let it finish or reconnect while the caller talks.
  • Turn compression on from the start. Without it, an audio-only session ends at 15 minutes no matter how well you resume.

On OpenAI Realtime, the 60-minute cap rarely matters for support calls. It does matter for long intake or dispatch calls. Plan a summarize-and-restart path for those.

Function calling on a live call

Tools behave differently on each model, and the differences surface under interruption.

OpenAI Realtime. The model emits a function call item. You run the function, send a `function_call_output` item with `conversation.item.create`, and then `response.create`. If the caller interrupts while your tool runs, the result still arrives later. Decide whether the call is still relevant before you trigger a new response.

Gemini Live. On `gemini-3.8-live`, asynchronous `NON_BLOCKING` calling is the default, per the model page. Set `behavior: BLOCKING` to restore the old pause-until-result mode. For non-blocking calls, each response sets `scheduling`: `INTERRUPT` speaks the result at once, `WHEN_IDLE` waits until the model finishes, and `SILENT` stores it for later. The Live API has no automatic tool handling, so you must answer every call with `send_tool_response`. The tools guide has the details. On interruption, the server discards pending calls and sends their IDs.

GPT-Live. Reasoning and tools are delegated to a backend. See our GPT-Live build guide for the delegation event flow.

The phone-specific rule is the same for all three. A canceled call is not a reverted side effect. If Gemini cancels a `book_appointment` call ID after your backend already wrote the booking, the booking stays. Make side-effecting tools idempotent with a key tied to the call and the confirmed slot. Treat "canceled" as "do not speak the result," not "undo."

Framework plugins have their own tool bugs. One example is LiveKit agents-js issue 2593, where Gemini `NON_BLOCKING` tool results were lost. Pipecat issue 5997 reports that `GeminiLiveLLMService` connection failures did not mark the service unusable. Pin versions and test tools under interruption after every upgrade. Our guide to tool call accuracy explains how to score these.

Find out what your phone agent does when callers interrupt
Truncation drift, echo self-interrupts, and duplicate tool calls only show up on real phone audio, which is exactly where independent evaluation tests your agent.
Book a demo

DTMF, transfer, and hangup from each path

Call control is where the model API ends and the phone network begins. The bidirectional Twilio stream has one hard rule. Twilio's docs say "the only way to stop a Stream is to end the call," or move the call to new TwiML.

ActionNative SIP (Realtime)Native SIP (GPT-Live)Twilio media streamLiveKit or Pipecat
Read keypad`input_audio_buffer.dtmf_event_received``transport.dtmf.received``dtmf` message with `dtmf.digit`Framework DTMF events
Transfer to a human`POST /v1/realtime/calls/{id}/refer``POST /v1/live/sessions/{id}/refer`Update the Call with new TwiML containing a DialSIP REFER or a room handoff
Hang up`POST /v1/realtime/calls/{id}/hangup``POST /v1/live/sessions/{id}/hangup`Update the Call status to `completed`Delete the room or end the session
Goodbye clipping riskLow if you wait for playoutMedium; no per-response done eventHigh; wait for the final mark echoDepends on plugin

Three practices prevent the common failures:

  • Expose transfer and hangup as tools. Have the tool return immediately. Execute the REFER or Call update only after the final goodbye audio has played. On a bridge, "played" means the last mark came back.
  • Do not trust the model to read digits from audio. Keypad tones on a G.711 stream may or may not survive. Use the out-of-band DTMF event and inject the digits as text. Our DTMF testing guide lists the cases to cover.
  • REFER needs carrier support. OpenAI returns 200 once the REFER is relayed to your SIP provider. The provider still has to honor it. Test transfers against your actual trunk, not a softphone.

For a full treatment of ending and transferring calls inside frameworks, see ending and transferring calls in LiveKit and Pipecat.

Decision table: which path should you ship?

CriterionNative SIPMedia-stream bridgeLiveKit or Pipecat plugin
ModelsOpenAI Realtime, GPT-LiveAny WebSocket modelAny supported plugin
Gemini LiveNot availableYes, with resamplingYes
Barge-in correctnessProvider handles playoutYou build the ledgerFramework handles, verify
Audio inspection or recording in your codeNo; audio bypasses youFull accessFull access
Mixing in a separate TTS or STTNoYesYes
Extra network hopsNoneTwo WebSocketsSFU or transport plus model socket
Ops burdenWebhooks and sidebandHighest: you run DSP and stateMedium: framework upgrades
Best forOpenAI-only, inbound, fast launchCustom control, Gemini, multi-modelTeams already on a framework

A short rule of thumb follows. If you are OpenAI-only and inbound-first, start with native SIP. If you need Gemini or must record and inspect audio in your own code, use a bridge or a framework. If you already run LiveKit or Pipecat for cascaded agents, use the plugin and keep one call path. Our LiveKit vs Pipecat telephony comparison helps with that choice.

Latency budget per architecture

Speech-to-speech models remove the STT and TTS hops, but the phone path adds its own. The table below is an illustrative budget, not a measurement. It assumes your bridge runs in the same cloud region as the carrier media edge and the model endpoint. Replace each row with your own numbers.

Hop (one turn, caller stops to first audio heard)Native SIPBridgeFramework
Carrier ingress and PSTN40 to 100 ms40 to 100 ms40 to 100 ms
Carrier to your server (WebSocket)010 to 40 ms10 to 40 ms (SIP to SFU)
Your server to model0 (carrier sends direct)20 to 60 ms20 to 60 ms
Turn detection waitVAD silence settingSameSame, or framework turn model
Model time to first audioModel-dependentModel-dependentModel-dependent
Return path and carrier playout40 to 100 ms70 to 200 ms70 to 200 ms
Resampling (polyphase)00 to 5 ms0 to 5 ms

The turn-detection wait is usually the biggest controllable term. LiveKit's OpenAI plugin example sets `server_vad` with `silence_duration_ms` of 500 ms. That wait happens before the model even starts. Tuning it is a trade between speed and cutting callers off. That trade is the subject of our Pipecat latency measurement guide.

Measure the real number from a dual-channel call recording. Mark the end of caller speech on one channel and the first agent audio on the other. Never use the model's internal timestamps. They miss the carrier buffer, which is the part the caller feels.

What it costs per minute, including the carrier

Here are the published rates as of October 4, 2026.

ItemRateSource
`gpt-realtime-2.1` audio input$32 per 1M tokens ($0.40 cached)Model page
`gpt-realtime-2.1` audio output$64 per 1M tokensModel page
Realtime audio tokenization1 token per 100 ms in, 1 per 50 ms outCosts guide
`gpt-live-1`$0.05 per minute, billed per second, plus backendModel page
`gemini-3.8-live` audio$0.005 per min in, $0.018 per min outGemini pricing
Twilio inbound, local number$0.0085 per minTwilio voice pricing
Twilio Media Streams$0.0044 per minTwilio voice pricing
Twilio Elastic SIP origination, local$0.0034 per minTwilio SIP pricing

The Realtime math is the tricky one. Fresh caller audio costs 600 tokens per minute, or $0.0192. Agent audio costs 1,200 tokens per minute, or $0.0768. The costs guide also notes that "the entire conversation is sent to the model for each Response." Later turns re-read all earlier audio, at the cached rate when the cache hits and the full rate when it misses.

Worked example (illustrative assumptions). Take a 4-minute inbound call with 10 turns. The caller speaks for 1.4 minutes and the agent for 1.6 minutes. Instructions and tools add 1,200 text tokens per turn.

  • Realtime model cost: about $0.16 per call with a perfect cache, and about $0.38 if half of the re-read history misses the cache. Output audio is $0.12 of that.
  • Realtime via native SIP: add 4 x $0.0034 = $0.014 for SIP origination. Total is $0.17 to $0.39, or $0.04 to $0.10 per minute.
  • Realtime via a Twilio media stream: add 4 x ($0.0085 + $0.0044) = $0.052. Total is $0.21 to $0.43.
  • GPT-Live via SIP: 4 x $0.05 = $0.20, plus $0.014 carrier, plus backend tokens. With an assumed $0.03 backend, the total is about $0.24.
  • Gemini 3.8 Live via a Twilio stream: input is $0.007 to $0.02, depending on whether only speech or all streamed audio bills. Output is 1.6 x $0.018 = $0.029. Twilio adds $0.052. The total is about $0.09, before bridge compute.

Two takeaways are not visible on any single pricing page. On Gemini, the carrier leg can cost more than the model, so a SIP trunk instead of Media Streams changes your unit economics. On Realtime, cache hit rate is the cost lever. Changing instructions or tools mid-call busts the cache from that point forward, which the costs guide warns about. Verify your own token counts from `response.done` usage on Realtime and from usage metadata on Gemini before you budget. Then convert to cost per resolved call, which is the number that matters.

Stacked cost per four-minute inbound call for Realtime over SIP, Realtime over a Twilio media stream, GPT-Live over SIP, and Gemini Live over a Twilio media stream, split into model and carrier cost

What the research says about speech-to-speech models on calls

Vendor docs tell you how to connect. Three research results tell you what to test once you have.

Turn-taking needs its own benchmark. Full-Duplex-Bench (Lin et al., ASRU 2025) evaluates spoken dialogue models on four behaviors: pause handling, backchanneling, smooth turn-taking, and user interruption. It uses automatic metrics such as how often the model takes over the turn when it should not. The practical lesson is that "accuracy" on a transcript says nothing about these behaviors. Each one needs scripted audio with known timing.

Models interrupt too aggressively and rarely backchannel. The Talking Turns study (Arora et al., ICLR 2025) used a turn-taking model trained on human conversations as a judge. It found that spoken dialogue systems "sometimes do not understand when to speak up, can interrupt too aggressively and rarely backchannel." On the phone, aggressive interruption shows up as the agent talking over callers who pause to find an account number. Your test set should include mid-sentence pauses of one to three seconds.

Degraded audio produces confident, invented content. The WildASR benchmark (2026) simulated phone channels with G.711 mu-law and GSM codecs, among other conditions. It found severe, uneven degradation that "does not transfer across languages or conditions." Models also hallucinated "plausible but unspoken content under partial or degraded inputs." That study measured ASR systems, but speech-to-speech models face the same 8 kHz input. Test your agent on codec-degraded audio, not on clean recordings.

How to test a realtime model on a phone line before launch

This is the protocol we recommend for an in-house team without an eval group. It assumes one path is built and a second is a candidate. Our GPT-Live testing guide covers model-level checks. This one targets the phone path.

1. Record a golden corpus over the real path. Place scripted calls through your actual trunk or bridge. Record dual-channel at the carrier, not in your app, so the carrier buffer shows up.

2. Build 8 scenarios. Use a clean happy path, a mid-sentence barge-in, a barge-in during a tool call, a 2-second thinking pause, DTMF entry, a transfer request, a goodbye and hangup, and a 14-minute call.

3. Cross them with 4 conditions. Use clean audio, babble noise at 10 dB SNR, 2 percent packet loss with jitter, and speakerphone echo. That gives 32 cells per path.

4. Score each cell with fixed metrics. Use the thresholds in the matrix below. They are starting targets for you to adjust, not industry standards.

5. Check context against audio after every barge-in. Compare what the model later says it told the caller with what the recording shows the caller heard. Any reference to unheard content is a truncation failure.

6. Count side effects in your backend, not in transcripts. A booking that exists twice is a failure even if the call sounded fine.

7. Run enough calls to see rare failures. If you see zero failures in n calls, the 95 percent upper bound on the true rate is about 3/n. To claim duplicate tool execution under 1 percent, you need about 300 clean interruption calls.

8. Re-run the full matrix on every model, plugin, or carrier change. Model snapshots and framework releases change interruption behavior without changing your code.

MetricDefinitionStarting pass threshold
Barge-in stop timeCaller speech onset to agent audio stopping, from the recordingp95 at or under 400 ms
Truncation fidelityShare of barge-ins where later turns reference only heard content98 percent or more
Self-interrupt rateAgent turns cut by its own echo, per 100 agent turnsUnder 1
False takeover rateAgent starts talking during a caller pause under 2 sUnder 5 percent
Duplicate side effectsTool writes executed twice per 100 tool calls0
Long-call survival14-minute calls finishing without a silence over 2 s100 percent
DTMF captureKeypresses delivered to the agent as digits100 percent
Goodbye completenessCalls where the final sentence plays before hangup98 percent or more

For a rate like truncation fidelity, size the sample with n = z^2 x p(1 - p) / E^2. To measure a 2 percent failure rate within plus or minus 2 points at 95 percent confidence, n = 3.84 x 0.02 x 0.98 / 0.0004, which is about 188 barge-in events. Scripted barge-ins make that affordable. Manual test calls do not. If you would rather have this matrix run by an outside team, as a pre-launch audit or a bake-off between Realtime, GPT-Live, and Gemini on your own scripts, that is the work Evalgent does.

The failure taxonomy

Use this table to label every failed call. Labels turn a pile of bad recordings into a ranked fix list.

LayerFailureTypical causePath most exposed
SignalingCall accepted twice or neverWebhook retries not deduplicated; competing accept handlersNative SIP
SignalingDuplicate outbound callAuto-retry after ambiguous timeoutGPT-Live outbound SIP
MediaOne-way audioSRTP media CIDRs blocked by firewallNative SIP
MediaStatic or noise burstsPCM16 sent where mu-law expected, or the reverseBridge
MediaMetallic sibilantsDecimating 24 kHz to 8 kHz without a low-pass filterBridge (Gemini)
Turn-takingAgent talks over caller pausesVAD silence too short; model behaviorAll
Turn-takingAgent interrupts itselfLine echo reflected into the inbound streamBridge, framework
ContextAgent references unheard wordsWrong `audio_end_ms`; Gemini keeps sent-but-unplayed audioBridge
SessionSilence at about 10 minutesGemini connection reset without resumptionBridge, framework
ToolsBooking written twiceNon-idempotent tools under cancel or retryAll
ToolsResult never spokenLost non-blocking results in a pluginFramework
Call controlGoodbye clippedHangup issued before final playoutAll
CostBill grows with call lengthContext re-read each turn; cache bustedRealtime

Where independent evaluation fits

Most failures in this taxonomy are invisible in the model vendor's dashboard. The model sees its own events, not the carrier buffer, the echo, or your backend's duplicate writes. They show up only when you test the whole phone path from the outside and score the recording, not the event log. That is the job of an independent evaluator such as Evalgent.

That helps at three points: before launch, when choosing between Realtime, GPT-Live, and Gemini on your own call scripts, and after every model or framework upgrade. If you want a structured outside review before go-live, see our third-party voice agent audit guide. For carrier choices, see the best telephony for voice agents and SIP vs WebRTC.

Frequently asked questions

Does the OpenAI Realtime API support SIP?

Yes. Point your SIP trunk at `sip:$PROJECT_ID@sip.api.openai.com;transport=tls`. OpenAI sends a `realtime.call.incoming` webhook. You accept with `POST /v1/realtime/calls/{call_id}/accept`, then attach a WebSocket using the `call_id` to handle events, tools, transfers via `/refer`, and hangups via `/hangup`. GPT-Live has a parallel flow under `/v1/live/sessions`.

Should I use OpenAI's SIP endpoint or Twilio Media Streams?

Use native SIP if you only need OpenAI models and don't need to process audio in your own code. OpenAI then handles playout and truncation. Use a Media Streams bridge if you need Gemini, recording, custom DSP, or several models on one number. The bridge costs more per minute on Twilio and you own barge-in correctness.

How do I compute audio_end_ms for conversation.item.truncate?

Use audio the caller actually heard, not wall-clock time. Name each Twilio mark with the cumulative milliseconds sent. When a mark echoes, that much audio has played. Freeze the count before sending `clear`, because Twilio echoes cleared marks too. Never exceed the audio sent, or the server rejects the truncate.

Does the Gemini Live API support SIP or Twilio?

Google's Gemini API docs list no SIP endpoint as of October 2026. Connect phone calls through a Twilio or Telnyx media-stream bridge, or through a framework like LiveKit or Pipecat. Decode mu-law and send PCM16 with the sample rate in the MIME type. Downsample Gemini's 24 kHz output to 8 kHz with a proper low-pass filter.

How long can a Gemini Live API session last on a phone call?

Audio-only sessions stop at 15 minutes unless you enable context window compression. Separately, connections reset after about 10 minutes, with a GoAway message first. Enable session resumption, store the latest handle, and reconnect. Handles stay valid for 2 hours. Buffer caller audio during the reconnect so no words are lost.

What happens to Gemini Live's context when a caller interrupts?

Gemini cancels and discards the ongoing generation and sends `serverContent.interrupted`. Everything already sent to your client stays in history. On a phone bridge, that includes audio still queued in the carrier's buffer, so the model may think the caller heard words they missed. Keep the carrier buffer shallow to limit this.

How do I transfer or hang up a call from a realtime voice agent?

On native SIP, call the `refer` or `hangup` endpoint for the call or session. On a Twilio bridge, update the Call with new TwiML containing a Dial to transfer, or set its status to `completed` to hang up. Wait for the final goodbye audio to finish playing first.

How much does an OpenAI Realtime phone call cost per minute?

It depends on talk ratio and cache hits. In our illustrative 4-minute example, `gpt-realtime-2.1` cost $0.16 to $0.38 per call. Native SIP added about $0.014 in Twilio origination, and a Media Streams bridge about $0.052. That works out to roughly $0.04 to $0.11 per minute. Check `response.done` usage on your own calls.

The bottom line

Pick native SIP when you are OpenAI-only and want the provider to own playout, and pick a bridge or a framework when you need Gemini, your own audio, or several models. Whichever path you choose, test barge-in, truncation, long calls, and side effects over the real phone line before callers find the bugs for you.

Related Articles