Evalgent
Back to Blog
Voice AI Testing

LiveKit + Twilio Outbound Calling: SIP Trunk Setup and What to Test Before Go-Live

Deepesh Jayal
25 min read
LiveKit + Twilio Outbound Calling: SIP Trunk Setup and What to Test Before Go-Live
On this page

Getting the first outbound call working takes about an hour. You create a Twilio trunk, paste the termination URI into a LiveKit outbound trunk, call `CreateSIPParticipant`, and your phone rings. The trouble starts in week two. Calls show up as "Spam Likely." Your agent talks over voicemail greetings. Unanswered calls leave agent jobs running. Twilio throttles you at one call per second, and you find out from your error logs.

This guide follows the call one leg at a time. It covers the SIP code each leg can return, what LiveKit does with that code, and what it costs. It ends with a go-live test matrix and pass criteria. For how LiveKit telephony compares with Pipecat, see our LiveKit vs Pipecat telephony comparison. This post covers outbound calling only, in more depth.

1 CPS
Twilio default termination calls per second, per trunk, per region (Twilio CPS docs)
80 s
Upper limit for ringing_timeout on CreateSIPParticipant (LiveKit SIP API)
$0.0100
Twilio Elastic SIP termination per minute, 48 US states (Twilio pricing, Aug 2026)
3% vs 35%
Share of unknown-number calls answered with and without a spam warning (Sherman et al., NDSS 2020)

The outbound call path, leg by leg

An outbound call is not one connection. It is five hops, run by three companies. Each hop can fail in its own way.

Ladder diagram of a LiveKit Twilio outbound call showing agent dispatch, CreateSIPParticipant, SIP INVITE with digest challenge, 180 or 183 ringing, 200 OK, ACK, and G.711 RTP audio to the PSTN

Leg 1: your backend to LiveKit (dispatch). Your scheduler calls the Agent Dispatch API with an `agent_name` and JSON metadata holding the phone number. The agent dispatch docs say LiveKit targets a max dispatch time under 150 ms. The room is created if it doesn't exist. An agent server process accepts the job.

Leg 2: the agent job to the LiveKit SIP service. Inside the job, your code calls `CreateSIPParticipant`. This request needs the SIP `call` grant. LiveKit validates it and picks a SIP node. Per the region pinning docs, outbound calls start in the region where the API call was made. If you set `destination_country`, they start from a server in that country instead.

Leg 3: LiveKit SIP to Twilio's termination URI. LiveKit sends an `INVITE` to `.pstn.twilio.com`. The SDP offer rides in that first INVITE. LiveKit's codec reference says outbound calls "always include an early offer." The default offer is PCMU, PCMA, and G.722. If you use credential auth, Twilio replies `407 Proxy Authentication Required`. LiveKit resends the INVITE with an `Authorization` header. The SIP handshake reference says this challenge "is not a failure." Expect to see it in every PCAP.

Leg 4: Twilio to the PSTN and the callee's carrier. Twilio routes the call onward. Provisional responses come back: `100 Trying`, then `180 Ringing` or `183 Session Progress`. When a mobile user answers, their carrier sends an SS7 Answer Message, which turns into a `200 OK` by the time it reaches LiveKit.

Leg 5: media. After `200 OK` and `ACK`, RTP flows. On a US PSTN call this is almost always G.711 μ-law (PCMU) at 8 kHz, about 64 kbps before overhead.

Where the audio gets resampled

People lose quality here without noticing. The phone network carries 8 kHz narrowband audio. Inside the LiveKit room, the callee is a normal participant with a WebRTC audio track. The SIP service converts between the two. Your agent's STT plugin then resamples room audio to whatever rate the STT model expects, often 16 kHz. Upsampling does not bring back the missing frequencies, though. The STT still hears a band that tops out near 4 kHz.

The return path does the same in reverse. Your TTS may render at 24 kHz, the room carries it, and the SIP service squeezes it back into 8 kHz PCMU. Sibilants and plosives that sounded crisp in the browser get dull.

How much does narrowband hurt recognition? The best controlled number we found is older. Bauer et al. (EUSIPCO 2014) trained matched recognizers on about 70 hours of German speech. The wideband system beat the narrowband one by 6.9% relative WER (36.83% vs 39.58%). The paper also cites a 20% relative WER reduction from doubling the sample rate from 8 to 16 kHz in earlier work. Those were HMM-era systems, so read the numbers as direction, not as a prediction for your model. Either way, a browser demo tells you little about phone accuracy. Test at 8 kHz. Our SIP vs WebRTC comparison covers why the two transports behave so differently.

You can offer wideband. LiveKit supports G.722 by default and AMR-WB as an opt-in. Whether you get it depends on every hop to the handset. Most PSTN calls to US mobiles still settle on PCMU. Plan for that.

LegProtocolTypical failureWhere you see it
Backend to LiveKitAgent Dispatch API`agent_name` mismatch, no agent server runningJob never starts, no SIP call logged
Agent to SIP serviceTwirp `CreateSIPParticipant`Wrong trunk ID (404 "object cannot be found")`SipCallError` or Twirp error in agent logs
SIP to TwilioSIP over UDP/TCP/TLS403 bad credentials, 400 non-E.164, 503 wrong addressLiveKit Telephony dashboard, Twilio debugger
Twilio to PSTNSIP to SS7486 busy, 480/408 no answer, 404 bad number`sip_status_code` on `SipCallError`
MediaRTP, G.711One-way audio, media timeout, garbled audioPCAP, RTP stream analysis

Setting up the Twilio side: Elastic SIP termination

On Twilio, outbound is called termination. Inbound is origination. Mixing the two up is the most common setup mistake.

1. Create an Elastic SIP trunk. The domain name must end in `pstn.twilio.com`, per the LiveKit Twilio quickstart.

2. On the Termination tab, set a termination SIP URI such as `acme-outbound.pstn.twilio.com`.

3. Attach authentication.

4. Associate at least one Twilio number with the trunk. It becomes your caller ID.

twilio api trunking v1 trunks create \
  --friendly-name "Acme outbound" \
  --domain-name "acme-outbound.pstn.twilio.com"

twilio api trunking v1 trunks phone-numbers create \
  --trunk-sid <twilio_trunk_sid> \
  --phone-number-sid <twilio_phone_number_sid>

Credentials vs IP access control

The Twilio SIP trunking docs say: "You must configure a minimum of either an ACL or credential authentication. If you configure both, then both ACLs and credentials are enforced."

  • Credential list. Pick a username and password in Twilio. Use the same pair on your LiveKit outbound trunk. This works from any LiveKit region, which is why LiveKit's quickstart uses it.
  • IP access control list. Allow only LiveKit's IPs. According to the outbound trunk docs, LiveKit Cloud publishes static IP ranges for Canada, the EU, India, Japan, and the US. Outside those regions there are no static ranges, so use credentials.

Using both is the strictest option. If you do, check that every region your agents dial from is on the ACL. When Twilio rejects you, its debugger logs error 32201 ("Source IP address not in ACL") or 32202 ("Bad user credentials"). LiveKit's troubleshooting guide maps these to a `403 Forbidden`.

Number format and caller ID rules

Twilio is strict about format. Numbers must be E.164 with the leading `+`. Otherwise "the call will be rejected with a SIP `400 Bad Request` response." Normalize numbers when you import contacts, not when you dial. A CSV with `(415) 555-0100` will fail at runtime and nowhere else.

Caller ID has a rule too. You must use "a Caller ID Number that either corresponds to a Twilio DID on your account or a Caller ID Number that has been verified." If you dial from a number Twilio doesn't recognize, the call fails before any phone rings.

International destinations are controlled by Twilio's Geo Permissions, which apply to Elastic SIP Trunking. Blocked calls log error 32205, "Geo Permission configuration is not permitting call." Trial accounts can only reach low-risk destinations.

Creating the LiveKit outbound trunk

The LiveKit outbound trunk holds the Twilio address, credentials, and caller ID numbers. Create it once and reuse it. The docs warn that "creating a new trunk for each call bypasses this caching and can degrade reliability at scale."

# Simplified. Field names match the LiveKit SIP API reference.
import asyncio, os
from livekit import api
from livekit.protocol.sip import CreateSIPOutboundTrunkRequest, SIPOutboundTrunkInfo

async def main():
    lkapi = api.LiveKitAPI()  # reads LIVEKIT_URL / API_KEY / API_SECRET
    trunk = SIPOutboundTrunkInfo(
        name="acme-twilio-outbound",
        address="acme-outbound.pstn.twilio.com",  # hostname only, no "sip:"
        numbers=["+15105550100"],                  # E.164, owned by Twilio account
        auth_username=os.environ["SIP_AUTH_USERNAME"],
        auth_password=os.environ["SIP_AUTH_PASSWORD"],
        destination_country="US",                  # region pinning for outbound
    )
    info = await lkapi.sip.create_sip_outbound_trunk(
        CreateSIPOutboundTrunkRequest(trunk=trunk)
    )
    print(info.sip_trunk_id)  # ST_xxxx, store it in config
    await lkapi.aclose()

asyncio.run(main())

Three details matter in production:

  • `address` is a hostname, not a URI. The SIP API reference says it "shouldn't contain the `sip:` protocol."
  • One trunk, many caller IDs. Set `numbers` to `["*"]` and pass `sip_number` on each `CreateSIPParticipant` call. Use this to rotate local-presence numbers without creating new trunks.
  • Inline config for multi-tenant apps. If each customer brings their own Twilio account, pass a `SIPOutboundConfig` inline on each call with `sip_number` set. This is the documented pattern for "a separate SIP provider per customer."

Dispatching the agent and placing the call

LiveKit supports two outbound patterns. Your backend can create the SIP participant and dispatch the agent to that room. Or the agent can dial from inside its own job. Most teams use the second. The agent owns the call from the first ring, so error handling stays in one place.

Explicit dispatch is required. Per the dispatch docs, "with `agent_name` set, the agent is only assigned to rooms when explicitly dispatched." Here is the trigger from your scheduler:

# Illustrative scheduler-side dispatch
import json, uuid
from livekit import api

async def place_outbound(phone_e164: str, contact_id: str):
    async with api.LiveKitAPI() as lkapi:
        await lkapi.agent_dispatch.create_dispatch(
            api.CreateAgentDispatchRequest(
                agent_name="outbound-reminder-agent",
                room=f"ob-{contact_id}-{uuid.uuid4().hex[:8]}",  # unique per attempt
                metadata=json.dumps({"phone_number": phone_e164, "contact_id": contact_id}),
            )
        )

Use a unique room per attempt. If you reuse a room name across retries, a late `participant_disconnected` event from attempt one can tear down attempt two.

Now the agent entrypoint. This version adds what the quickstart leaves out: a ring timeout, a duration cap, outcome mapping, and explicit shutdown.

# Simplified agent entrypoint. Verify against your livekit-agents version.
import json
from google.protobuf.duration_pb2 import Duration
from livekit import agents, api

TRUNK_ID = "ST_xxxx"

@server.rtc_session(agent_name="outbound-reminder-agent")
async def entrypoint(ctx: agents.JobContext):
    info = json.loads(ctx.job.metadata)
    phone = info["phone_number"]
    try:
        await ctx.api.sip.create_sip_participant(api.CreateSIPParticipantRequest(
            room_name=ctx.room.name,
            sip_trunk_id=TRUNK_ID,
            sip_call_to=phone,
            participant_identity=phone,
            wait_until_answered=True,                 # block until 200 OK or failure
            ringing_timeout=Duration(seconds=30),     # documented max is 80 s
            max_call_duration=Duration(seconds=600),  # hard cap on runaway calls
        ))
    except api.SipCallError as e:
        # 486/603 -> rejected, 408/480 -> no answer, 5xx -> trunk failure
        record_attempt(info["contact_id"], e.sip_status_code, e.sip_status)
        ctx.shutdown()   # required: some failure reasons do NOT auto-close the job
        return

    participant = await ctx.wait_for_participant(identity=phone)
    # Start the AgentSession here, AFTER answer. Do not greet first on outbound.

The docs are clear about when to start the session: "Call `session.start()` after the callee picks up. If the session starts while the call is still ringing, the initial greeting plays before the callee joins the room." A recurring GitHub complaint, the bot starting to talk before pickup, comes down to this mistake.

Answer vs ringing vs busy vs no-answer

`wait_until_answered=True` turns the SIP outcome into a Python result. A `200 OK` returns. Anything final and non-2xx raises `SipCallError`, which carries the carrier's code. This is how each code maps, based on LiveKit's outbound calls guide and SIP participant reference:

SIP responseMeaningLiveKit result`disconnect_reason`Auto-closes session?What your code should do
`180 Ringing`Phone is ringing, no mediaStill waiting, `sip.callStatus = dialing`n/an/aNothing. Ring clock runs
`183 Session Progress`Ringing with early mediaStill waitingn/an/aDon't treat carrier audio as the callee
`200 OK`Answered (human or machine)Returns, `sip.callStatus = active`n/an/aRun AMD before speaking
`486 Busy Here` / `603 Decline`Busy or rejected`SipCallError``USER_REJECTED`YesRetry later with backoff
`408 Request Timeout` / `480 Temporarily Unavailable`No answer or unreachable`SipCallError``USER_UNAVAILABLE`NoCall `ctx.shutdown()`, schedule retry
`404 Not Found`Bad number or bad trunk ID`SipCallError`variesCheckMark number invalid, stop retrying
`5xx`Trunk or protocol failure`SipCallError``SIP_TRUNK_FAILURE`NoCall `ctx.shutdown()`, alert on rate

Look at the "Auto-closes" column. LiveKit's docs say `AgentSession` "automatically closes the session when a SIP participant disconnects with `USER_REJECTED`. If the disconnect reason is `USER_UNAVAILABLE` or `SIP_TRUNK_FAILURE`, you must explicitly call `ctx.shutdown()` to release the job." No-answer is your most common outcome. Skip the shutdown and every no-answer leaves a job running. That job holds a slot on an agent server that should be dialing the next number.

Voicemail is the other trap. Voicemail systems answer with `200 OK`. To LiveKit, that is a successful call. The docs put it plainly: "Voicemail is not a failure." Telling a person from a machine is a separate step.

Answering machine detection and the first-word timer

LiveKit's answering machine detection runs once, on the first thing the callee says. It returns one of five categories: `human`, `machine-ivr`, `machine-vm`, `machine-unavailable`, or `uncertain`. Agent speech stays paused until the result arrives. Two paths run at once: a fast heuristic for short greetings and an LLM classifier for longer ones.

The defaults matter for timing:

  • `human_speech_threshold`: 2.5 s. Speech shorter than this takes the fast "human" path.
  • `human_silence_threshold`: 0.5 s. Silence needed after a short greeting before AMD decides `human`.
  • `machine_silence_threshold`: 1.5 s. Silence after machine-like speech before a verdict.
  • `no_speech_threshold`: 10 s. With no speech at all, AMD settles on `uncertain`.
  • `wait_until_finished`: `True`. AMD waits for the greeting to end, so a long greeting can run past `timeout`.

For comparison, Altwlkany et al. (arXiv 2024) built a small streaming GRU classifier on about 4,200 real call recordings. It reached 96.67% test accuracy, and 98.10% with a silence detector added. Inference took 31.63 ms per frame on CPU. The paper also cites vendor AMD running about 4 seconds on average with "above 90% accuracy in the US." Those vendor figures are claims the authors quote, not measurements they made. The useful point: AMD always trades accuracy against time, and every second of silence after "Hello?" is a second the callee spends deciding whether to hang up.

Timeline from callee answer to the agent's first word, showing a short hello, AMD silence threshold, LLM time to first token, TTS first byte, and network hops, compared with the 208 ms average human response gap

Here is a worked example of the first-word gap on a human pickup. Every component time below is an assumption, not a measurement.

StepAssumed timeRunning gap after callee stops talking
Callee says "Hello?"0.6 s of speech0 ms
AMD `human_silence_threshold`500 ms500 ms
LLM time to first token (assume)400 ms900 ms
TTS first audio byte (assume)150 ms1,050 ms
Room, SIP bridge, carrier to handset (assume)150 ms1,200 ms

Compare that 1.2 s gap with ordinary conversation. Stivers et al. (PNAS 2009) measured question-response pairs in 10 languages. The mean response offset was 208 ms, and every language averaged within 500 ms. Your first turn is the one that most often decides whether someone hangs up, and it is roughly five times slower than a human reply.

Ways to shorten it:

  • Pre-render the opener. The first line ("Hi, this is Maya from Acme Dental about your appointment") rarely changes. Generate it ahead of time and play it with `session.say`, so there is no LLM wait on turn one.
  • Warm the pipeline before dialing. Open STT, LLM, and TTS connections while the phone rings. LiveKit's AMD example builds the detector around the session before `create_sip_participant`, so this fits the documented flow. Just don't generate speech until you get the AMD verdict.
  • Tune `human_silence_threshold` with data. Lowering it from 500 ms saves time but sends more machines down the human path. Measure both sides on recorded greetings before you change it.

For latency measurement in general, see our guide to time to first audio.

Ending the call: EndCallTool and hangup semantics

Outbound calls need a clean exit. The prebuilt EndCallTool is in beta for Python and Node.js. When the LLM calls `end_call`, four things happen in order. The agent says a final line (from `end_instructions`). The session shuts down after that line plays. The room is deleted if `delete_room` is `True` (the default). Then the job process exits.

from livekit.agents import Agent
from livekit.agents.beta.tools import EndCallTool

class ReminderAgent(Agent):
    def __init__(self):
        end_call = EndCallTool(
            extra_description="Only end the call after the appointment is confirmed, "
                              "rescheduled, or the person asks to stop.",
            delete_room=True,
            end_instructions="Thank them and say goodbye in one short sentence.",
        )
        super().__init__(instructions="You confirm dental appointments.",
                         tools=end_call.tools)

Keep `delete_room=True`. The outbound guide warns: "If the agent session ends but the room is not deleted, the user continues to hear silence until they hang up." Silence is billable. Twilio charges for connected minutes until someone sends a `BYE`.

Deleting the room shows up as `ROOM_DELETED` in `disconnect_reason`. When the callee hangs up, a clean `BYE` shows up as `CLIENT_INITIATED`. Log both. If you see many calls where the callee hangs up within 10 seconds of answering, look at your opener and your caller ID reputation, not at your prompt.

Set `max_call_duration` as a backstop. If an LLM loop never calls `end_call`, or the callee leaves the phone off the hook, the cap ends the call. Ten minutes is a sensible ceiling for most reminder and qualification flows. For call transfers rather than hangups, see our post on DTMF and IVR navigation testing.

Caller ID, STIR/SHAKEN, and spam labels

Outbound has a problem inbound doesn't. The callee decides whether to pick up based on what their screen shows. That screen depends on STIR/SHAKEN attestation and on carrier analytics.

Twilio's SHAKEN/STIR docs define the levels:

  • A: "the caller is known and has the right to use the phone number as the caller ID."
  • B: "the customer is known, it is unknown if they have the right to use the caller ID."
  • C: everything else, "including international calls."

To get A, your Twilio account needs an approved Business Profile and an approved SHAKEN/STIR Trust Product, with the number assigned to both. Per Twilio's onboarding docs, B is "the highest level of attestation possible if a customer is using non-Twilio phone numbers." If you bring numbers ported from elsewhere and send them through a Twilio trunk without porting them in, you are capped at B. On Elastic SIP Trunking, Twilio passes the result in the `X-Twilio-VerStat` and `Identity` headers. Check those headers in your test calls rather than assuming you got A.

Why it matters: Sherman et al. (NDSS 2020) ran a lab study with 34 participants and five incoming-call screen designs. With no warning, participants answered 35% of calls from unknown numbers. With a spam warning, that fell to 5% and 3%, depending on the design. An authenticated caller ID notice raised it to 42%. A spam label even suppressed answers from known numbers: 100% with no warning, 34% to 65% with one. The sample is small and the setting is a lab. Still, the direction is hard to miss. Your answer rate can drop by an order of magnitude before your agent says a word.

Attestation is only part of it. Carrier analytics engines label numbers based on behavior: volume per number, short call durations, and complaint rates. Twilio's Voice Integrity registers numbers with the analytics engines for T-Mobile, Verizon, and AT&T to "remediate spam labels." It requires an approved Business Profile with an EIN or DUNS number. Do this before launch. Removing a label after the fact is slower.

One more data point for your campaign design. Prasad et al. (USENIX Security 2020) ran a honeypot of up to 66,606 lines for 11 months. They found that most robocall campaigns "rarely reuse phone numbers." Carriers see number churn as a robocall signal. Rotating through dozens of fresh numbers to avoid labels can make your traffic look more like the campaigns carriers are trying to block.

Telnyx as the alternative trunk

LiveKit supports Twilio, Telnyx, Plivo, Wavix, Sinch, and didlogic with provider quickstarts. For a full carrier comparison, see our telephony provider roundup. For outbound specifically, these Telnyx differences are documented:

ItemTwilio Elastic SIPTelnyx
Outbound trunk `address``.pstn.twilio.com``sip.telnyx.com` or a regional signaling address
Common address mistakeUsing the origination URIAdding a subdomain, which returns `503` (LiveKit troubleshooting)
Number formatE.164 with `+`, else `400`Leading `+` assumes "Destination Number Format" is `+E.164`
SIP REFER transfersSupportedMust be enabled on your account (LiveKit troubleshooting)
US outbound price (listed)$0.0100/min, 48 states"Starting at $0.005 per minute" local
Toll-free outbound$0.0011/minListed as free

Prices are from the Twilio US SIP pricing page (marked current as of August 2026) and the Telnyx Elastic SIP pricing page. Telnyx's "starting at" price is a floor. Get a quote for your actual traffic. We didn't find Telnyx's billing increment on the pricing page, so ask for it.

Switching trunks in LiveKit means changing the `address`, credentials, and numbers. Your agent code doesn't change. That makes it practical to run a small share of calls through a second carrier and compare answer rates, SIP code distributions, and audio quality on the same scripts.

What an outbound minute costs

You pay three meters. Here are the list prices, verified October 2026:

  • Twilio termination: $0.0100/min to the 48 states. Partial minutes round up. Twilio's support article says calls "under 60 seconds are rounded up to the next full minute" on Elastic SIP Trunking. That article is on a legacy support site, so confirm against your own invoice.
  • LiveKit third-party SIP minutes: "Inbound and outbound minutes using a third-party SIP trunk." On the Ship plan, 5,000 are included, then $0.004/min. On Scale, 50,000 are included, then $0.003/min (LiveKit pricing).
  • LiveKit agent session minutes: $0.01/min after the plan's included minutes, on Ship and Scale.

On Ship overage, that is $0.024 per connected minute before STT, LLM, and TTS. The per-minute rate isn't the hard part. Rounding, ringing, and short calls are.

Worked example: 10,000 attempts

Assumed mix (illustrative, not measured): 30% human answers averaging 2.5 minutes. 25% voicemail, handled in 25 seconds. 45% no answer or busy. Average ring time: 12 s before a human answers, 15 s before voicemail, 25 s for no-answer (with `ringing_timeout` at 30 s). We assume Twilio doesn't charge for unanswered attempts. Check your own invoice.

Line itemCalculationResult
Connected minutes3,000 x 2.5 + 2,500 x 0.4178,542 min
Twilio billed minutes (round up)3,000 x 3 + 2,500 x 111,500 min
Twilio cost11,500 x $0.0100$115.00
LiveKit SIP minutes (no included minutes)8,542 x $0.004$34.17
Agent job time if billed from dispatch10,000 x 69.85 s / 6011,642 min
Agent session cost (upper bound)11,642 x $0.01$116.42
Total before models$265.59
Per human conversation$265.59 / 3,000$0.089

Two results surprise people. First, rounding adds 35% to Twilio billed minutes in this mix (11,500 vs 8,542). Every 25-second voicemail is billed as a full minute. Second, the agent job is alive during ringing. With the agent-dials pattern, the job starts at dispatch, not at answer. We couldn't confirm from LiveKit's docs whether agent session minutes accrue during ringing. Check your usage dashboard after a test batch. The table uses the worst case.

Calls per second and concurrency

Twilio's CPS docs set the default at "1 CPS per Trunk per Region." You can raise it to 5 in the console. Higher limits go through sales. Going over logs errors 32001 and 32012. At 1 CPS, the ceiling is 3,600 attempts per hour, per trunk, per region.

Little's law gives you agent concurrency: L = λ x W, where λ is the attempt rate and W is how long each attempt holds a job.

  • W = 0.45 x 25 s + 0.25 x (15 + 25) s + 0.30 x (12 + 150) s = 11.25 + 10 + 48.6 = 69.85 s
  • At λ = 1 attempt per second, L = about 70 concurrent agent jobs
  • Ringing seconds = 0.45 x 25 + 0.25 x 15 + 0.30 x 12 = 18.6 s, so 27% of agent capacity is spent waiting on ringing

Size your agent servers for that 70, plus headroom, before you raise CPS. Raising CPS to 5 without adding capacity is how teams see the concurrency failures that never appear in single-call tests.

Twilio's CPS page also lists conditions for getting higher limits. Average call duration must be over 30 s. No more than 10% of calls can be 12 s or shorter. Answer-seizure ratio must be over 70%. Outbound AI campaigns break these by design. AMD hangs up on full mailboxes in a few seconds. Cold lists rarely hit 70% answered. In the example mix, 55% of attempts are answered. Before you ask Twilio for more throughput, find out whether your traffic qualifies. And log short calls as their own metric.

Know your outbound failure rates before your callers do
Evalgent runs independent pre-launch audits of LiveKit outbound agents across real carriers, voicemail systems, and edge-case numbers, then scores every call against pass criteria you can act on.
Book a demo

Failure taxonomy for LiveKit + Twilio outbound calls

Each failure below either has its own SIP code or has a specific signature in the media or in agent logs. Group them by layer so you know which dashboard to open first.

Failure taxonomy grid for LiveKit Twilio outbound calls in four layers: SIP signaling before answer, media after answer, agent lifecycle, and reputation and compliance, with the code or signal for each failure
FailureCode or signalLikely causeFix
Digest challenge`401` / `407`Normal auth handshakeNone. Not a failure
Bad number format`400 Bad Request`Not E.164, missing `+`Normalize when you ingest contacts
Auth rejected`403`, Twilio 32201/32202Credential mismatch, IP not in ACLMatch the credential list. Add LiveKit static IPs
Wrong region`403 Domestic Anchored Terms Not Met`Call left the required countrySet `destination_country`
Trunk not found`404` "object cannot be found"Stale `ST_` ID in configLoad the trunk ID from config, check at startup
Number not in service`404` from destinationDisconnected numberMark invalid, suppress retries
Codec mismatch`488 Not Acceptable Here`No common codec in SDPKeep PCMU in the offer. Avoid `only_listed_codecs` without it
Wrong trunk address`503`Subdomain on Telnyx address, typoUse the provider's exact signaling host
Geo blockedTwilio 32205Destination country not allowedEnable it in Geo Permissions, or block it in your app first
ThrottledTwilio 32001/32012Over the CPS limitPace dials in your scheduler
One-way audioRTP only one directionNAT. Self-hosted SIP advertising a private IPSet `use_external_ip: true` when self-hosting
Media timeoutCall drops, "media timeout"No RTP for 30 s at start or 15 s mid-callRaise `media_timeout` (max 10 minutes)
Garbled audioStatic, robotic voiceRTP payload type mismatch, loss over 3%, jitter over 20 msPCAP, Wireshark RTP stream analysis
Greeting before pickupCallee hears half a sentenceSession started before `200 OK`Start the session after `wait_until_answered`
Talking to voicemailAgent pitches a mailboxNo AMD, or AMD misclassifiedAMD, and test misclassifications on recorded greetings
Leaked jobsAgent servers fill upNo `ctx.shutdown()` on 408/480/5xxShut down in every `SipCallError` branch
Dead air at hangupCallee hears silenceSession ended, room not deleted`EndCallTool(delete_room=True)`
Extension never dialedStuck at company IVR`dtmf` timing too fastAdd `w` pauses (0.5 s each) to the `dtmf` string

The quality thresholds come from LiveKit's troubleshooting guide. Packet loss under 1% is healthy, and over 3% "causes audible breakup." Mean jitter under 5 ms is healthy, and over 20 ms "causes choppy audio." One-way latency under 150 ms is healthy, and over 300 ms makes people talk over each other. That 150 ms figure matches ITU-T G.114. It says delays under 150 ms give "essentially transparent interactivity" and recommends staying under 400 ms for network planning.

When a call fails, LiveKit's guide suggests three questions before you download a PCAP. Did an INVITE get recorded? What SIP response came back? Did media flow after the handshake? Record all three for every call. Our post on what to log on every voice agent call covers the full field list, including `sip.callID`, `sip.twilio.callSid`, and `disconnect_reason`.

Outbound compliance, briefly

This isn't legal advice. These are the primary sources your counsel will ask about.

  • AI voices count as "artificial" under the TCPA. In FCC 24-17 (February 2024), the FCC ruled that the TCPA restriction on "artificial or prerecorded voice" covers "current AI technologies that generate human voices." Callers need prior express consent. Telemarketing calls need prior express written consent.
  • Calling hours. Under 47 CFR 64.1200(c)(1), telephone solicitations to residential subscribers can't be made "before the hour of 8 a.m. or after 9 p.m. (local time at the called party's location)." Enforce this in your scheduler using the callee's time zone, not your server's.
  • One-to-one consent. The FCC's one-to-one consent rule was vacated by the Eleventh Circuit in January 2025, before it took effect. The current eCFR text doesn't include it.

Store the consent record ID in the dispatch metadata. Then every call in your logs can be traced back to the consent behind it. For a full review, see our voice agent compliance audit guide.

How to test LiveKit + Twilio outbound calling before go-live

Telephony bugs only show up on real phone networks. Test on real numbers, on real carriers, at real volume. Here is a protocol you can run in two to three days.

1. Build a test number bank. Get at least one line on each of AT&T, Verizon, and T-Mobile, plus one landline or VoIP line. Add a number set to always-busy, one that never answers, one that goes straight to a full voicemail box, one disconnected number, and one IVR with an extension.

2. Check the signaling paths first. Make one call per scenario with `wait_until_answered=True`. Confirm the `sip_status_code` and `disconnect_reason` match the outcome table above. Confirm every failure branch calls `ctx.shutdown()` by checking active job counts after the batch.

3. Check caller ID and attestation on every carrier. Look at the callee's screen on each test line. Record what shows up: number, name, or a spam label. Capture `X-Twilio-VerStat` from the call. If you see anything other than A, fix the Trust Hub setup before launch.

4. Run AMD against recorded greetings. Collect at least 50 real voicemail greetings (carrier default, custom, short, long, non-English) and 50 human pickups ("Hello?", "Yeah?", silence, background noise). Score `human` vs `machine-vm` accuracy and the time from answer to verdict.

5. Measure the first-word gap. For human pickups, measure the time from the end of the callee's first utterance to the agent's first audio. Do this in the call recording, not the agent logs. Log p50 and p95.

6. Test the endings. Trigger `end_call` from the conversation, hang up from the callee side, and let `max_call_duration` expire. Each one should end the Twilio leg within a few seconds. Check the Twilio call duration against the LiveKit call record.

7. Load test at your target CPS. Run a batch at your planned dial rate for 30 minutes. Watch Twilio errors 32001/32012, agent server CPU, job counts, and the share of calls that drop into media timeout.

8. Re-run after every change. Prompt edits, model swaps, and SDK upgrades all change the first-word gap and AMD behavior. Treat this matrix as a regression suite, not a one-time launch checklist. Our LiveKit voice agent testing guide covers the conversation-level tests that sit on top of these telephony checks.

Go-live test matrix with pass criteria

ScenarioExpected signalCalls per carrierPass criteria
Human answers, mobile`200 OK`, AMD `human`1030/30 connect. First-word gap p95 under your target (example: 1.5 s)
Human answers, landline`200 OK`, AMD `human`1010/10 connect. Two-way audio confirmed
Busy`486` -> `USER_REJECTED`5Correct code. Job closes. Retry scheduled
No answer`408`/`480` -> `USER_UNAVAILABLE`5Job closes within ringing timeout + 5 s
Voicemail, room for a message`200 OK`, AMD `machine-vm`10Agent waits for the beep. Message under 20 s. Hangs up
Full mailboxAMD `machine-unavailable`5Hangs up without speaking
Disconnected number`404`3Number marked invalid. No retry
International, blockedTwilio 32205 or app-level block2Blocked before dialing, or refused cleanly
IVR with extension`sip.callStatus = automation` then `active`5Correct extension reached
Agent ends call`ROOM_DELETED`10Twilio leg ends within 3 s of goodbye
Callee hangs up mid-sentence`CLIENT_INITIATED`10Session closes. No orphaned job
Load at target CPSNo 32001/32012 errors1 batchZero throttling. Job count matches Little's law estimate ±20%

How many calls are enough? Use the rule of three. If you see zero failures in n independent trials, the 95% upper bound on the true failure rate is about 3/n. Thirty clean calls to mobile lines tells you the failure rate is probably under 10%, not that it's zero. To claim under 1%, you need about 300 clean calls. Plan the batch size around the risk you're willing to launch with.

Where independent evaluation fits

Most of this matrix is plumbing. Your team can script it with the LiveKit CLI and a few test lines. Three parts are harder to do from the inside.

The first is AMD and first-turn quality across greeting types. That takes a large, varied set of greetings and pickups, not the five your team recorded on their own phones. The second is scoring every production call, not a sample. Outbound failure modes like rising short-call rates, spam labels on one carrier, or AMD drift after a model change show up as trends across thousands of calls. The third is vendor bake-offs: Twilio vs Telnyx answer rates, or one TTS vs another at 8 kHz, scored by someone who doesn't sell either.

That is the work Evalgent does as an independent evaluator. We run pre-launch audits and regression checks for in-house LiveKit agents, and score production calls against your own pass criteria. For outbound-specific metrics, see our guide to outbound sales voice agent metrics.

Frequently asked questions

How do I make an outbound call with LiveKit and Twilio?

Create a Twilio Elastic SIP trunk with a termination URI and a credential list. Create a LiveKit outbound trunk with that address, the same credentials, and your Twilio number. Then dispatch an agent with the phone number in metadata and call `CreateSIPParticipant` with `wait_until_answered=True` from the agent's entrypoint.

Why does my LiveKit agent start talking before the callee picks up?

Your session starts and greets before the call is answered. Call `session.start()` only after `create_sip_participant` returns with `wait_until_answered=True`. On outbound calls, don't greet first. Let AMD classify the pickup, then play your opener. LiveKit's docs warn that early greetings get cut off or play into silence.

What does a SipCallError with code 486 or 480 mean?

486 Busy Here and 603 Decline mean the callee rejected the call. LiveKit maps them to `USER_REJECTED` and closes the session automatically. 408 and 480 mean no answer or unreachable, mapped to `USER_UNAVAILABLE`. That one does not auto-close. Call `ctx.shutdown()` yourself, or the job stays alive.

Does wait_until_answered detect voicemail?

No. Voicemail systems answer with `200 OK`, so `wait_until_answered` returns successfully. Use LiveKit's answering machine detection to classify the pickup as `human`, `machine-vm`, `machine-unavailable`, `machine-ivr`, or `uncertain`. Start AMD before you create the SIP participant, and branch on the result.

How do I hang up a LiveKit outbound call from the agent?

Add the prebuilt `EndCallTool` to your agent's tools and keep `delete_room=True`. When the LLM calls `end_call`, the agent says goodbye, the session shuts down, and the room is deleted, which ends the SIP leg. If the room isn't deleted, the callee hears silence and the carrier keeps billing.

Should I use Telnyx or Twilio with LiveKit for outbound calls?

Both work through the same outbound trunk API. Telnyx lists a lower starting outbound price. Twilio has documented SHAKEN/STIR attestation onboarding and Voice Integrity for spam label remediation. Since switching only changes the trunk address and credentials, route a small share of calls through each and compare answer rates on identical scripts.

How much does a LiveKit Twilio outbound call cost per minute?

At October 2026 list prices: Twilio termination is $0.0100/min to the 48 states. LiveKit third-party SIP minutes are $0.004/min on Ship after included minutes. Agent session minutes are $0.01/min. That totals about $0.024 per connected minute before STT, LLM, and TTS. Twilio rounds partial minutes up.

Why are my outbound calls showing as spam likely?

Usually low attestation, number behavior, or both. Get A-level attestation through a Twilio Business Profile and SHAKEN/STIR Trust Product. Register numbers through Voice Integrity. Avoid patterns carriers treat as robocalls: high volume per number, many very short calls, and constant number rotation. Check the label on each major carrier before launch.

The bottom line

A LiveKit Twilio outbound call is five legs, and most production failures come from three things the quickstart skips: SIP outcomes that don't auto-close the job, voicemail that looks like a successful answer, and caller ID reputation that decides whether anyone picks up. Build the outcome table, the cost math, and the go-live matrix into your launch checklist, and re-run them after every change.

Related Articles