Evalgent
Back to Blog
Voice AI Testing

Voice Agent Cost by Use Case: AI Voice Agent Cost Per Call and Per Month for 24 Use Cases at 10k–1M Minutes

Deepesh Jayal
19 min read
Voice Agent Cost by Use Case: AI Voice Agent Cost Per Call and Per Month for 24 Use Cases at 10k–1M Minutes
On this page

Most "how much does a voice agent cost" pages give you one per-minute number. That number is close to useless for planning. You budget calls, not minutes. And you get paid for resolved calls, not calls.

This post builds a cost model for a team that runs its own stack on LiveKit or Pipecat. It prices 24 use cases per call, per month at 10k, 100k and 1M minutes, and per resolved call. Every price is a public list rate I checked on October 5, 2026, with a link. Every workload number (call length, turns, prompt size, resolution rate) is a labeled assumption you should replace with your own. The Python calculator at the end lets you do that in a minute.

For the general per-minute pricing explainer, see AI voice agent cost. For people, maintenance and the rest of the bill, see voice agent total cost of ownership. This post is only about the runtime cost of a call.

$0.0065/min
Deepgram Flux English streaming, current list price (Deepgram)
0.1x
GPT-5.6 cached input price vs uncached input (OpenAI)
60 seconds
Minimum billed per call on Twilio and Telnyx voice (support docs)

The per-call cost formula

A cascaded voice agent call has five meters running at once. Each one bills in a different unit, which is why per-minute quotes hide so much.

Cost per call = STT + TTS + LLM + hosting + telephony, where:

  • STT = call minutes × STT rate. You stream the caller's audio for the whole call, silence included.
  • TTS = characters spoken by the agent ÷ 1,000 × TTS rate.
  • LLM = sum over every LLM request of (prompt tokens × input rate) + output tokens × output rate.
  • Hosting = call minutes × (agent session rate + observability rate).
  • Telephony = billed carrier minutes × carrier rate + call minutes × SIP bridge rate.

Here are the baseline rates. These are pay-as-you-go list prices, not negotiated ones.

ComponentBaseline choiceList price (Oct 5, 2026)Source
STTDeepgram Flux English, streaming$0.0065/min (promo; regular $0.0077)Deepgram pricing
LLMOpenAI GPT-5.6 Luna$0.20 input, $0.02 cached input, $1.20 output per 1M tokensOpenAI pricing
TTSDeepgram Aura-2$0.030 per 1,000 charactersDeepgram pricing
Agent hostingLiveKit Cloud agent session$0.01/minLiveKit pricing
ObservabilityLiveKit agent observability$0.01/minLiveKit pricing
SIP bridgeLiveKit third-party SIP minutes$0.004/min (Ship), $0.003/min (Scale)LiveKit pricing
Carrier, inboundTwilio Elastic SIP origination, US local$0.0034/minTwilio SIP pricing
Carrier, outboundTwilio Elastic SIP termination, lower 48$0.0100/minTwilio SIP pricing

GPT-5.6 has one more price you need. According to OpenAI's prompt caching guide, cache writes cost 1.25x the uncached input rate and cache reads cost 0.1x. The minimum cacheable prefix is 1,024 tokens. A cached prefix stays reusable for at least 30 minutes after its last write or reuse.

The workload assumptions behind every number

Prices are the easy half. The hard half is the workload: how long calls run, how much the agent talks, and how many tokens each turn drags through the model. These are my assumptions, with the reasoning behind each one.

Characters per spoken minute: 800. Cartesia's pricing FAQ says one minute of generated audio takes 750 to 800 characters (Cartesia pricing). That lines up with conversational speech research. Yuan, Liberman and Cieri measured 193 words per minute across 943,044 segments of the Fisher English telephone corpus. Turns averaged about 10 words, and assigned topics ranged from 152 to 170 words per minute (Interspeech 2006). TTS voices for agents usually speak slower than people on the phone, so 800 characters is a fair ceiling.

Agent talk share: 30–55% of the call. The agent talks more on reminders and outbound pitches, and less on intake calls and recruiting screens, where the caller narrates. TTS characters = talk share × call minutes × 800.

Four characters per token. That's the usual rule of thumb for English text. Output tokens per call = TTS characters ÷ 4.

30 tokens per caller turn. That's about 22 words. It matches the Fisher turn length above.

225 tokens per tool call, for the call JSON plus a compact result. Each tool call adds one more full-context LLM request. LiveKit's docs confirm that tool calls running after the first completion emit their own `LLMMetrics` events (LiveKit data hooks).

System prompt plus tool schemas: 1,200–6,000 tokens, depending on the use case. A menu, a policy summary or ten tool definitions add up quickly.

Telephony rounding: +0.5 billed minutes per call on average. Twilio rounds partial minutes up to the next full minute on Programmable Voice, Elastic SIP Trunking and Media Streams (Twilio support). Telnyx bills 60/60, so an 11-second call is billed as 60 seconds (Telnyx billing increments). When call lengths are spread out, rounding up adds about half a minute per call on average.

Cost per call and per month for 24 use cases

The table uses the baseline stack. Monthly totals include a LiveKit plan fee: $50 (Ship) at 10k minutes, and $500 (Scale) at 100k and 1M minutes. They leave out included minutes, which makes them slightly conservative. "Dir" means inbound or outbound. All call lengths, turns, prompt sizes and tool counts are illustrative assumptions.

Use caseDirAvg minAgent turnsPrompt tokensToolsCost/callCost/min10k min/mo100k min/mo1M min/mo
Customer support (general)in4.5124,0002$0.205$0.0455$505$5,054$46,037
Technical support / troubleshootingin8206,0004$0.364$0.0455$505$5,051$46,014
Appointment schedulingin393,0003$0.137$0.0458$508$5,076$46,260
Appointment remindersout1.241,5001$0.070$0.0582$632$6,324$58,739
Order status (WISMO)in262,5001$0.094$0.0471$521$5,214$47,639
Food orderingin4145,0002$0.178$0.0445$495$4,954$45,040
Restaurant reservationsin272,0002$0.092$0.0460$510$5,098$46,477
Billing inquiriesin4114,0002$0.182$0.0456$506$5,060$46,098
Collections (early-stage)out3.5104,5002$0.191$0.0545$595$5,947$54,970
Inbound lead qualificationin4113,0002$0.177$0.0443$493$4,932$44,818
Outbound sales / appt settingout393,5001$0.167$0.0558$608$6,083$56,326
SMB front-desk receptionistin262,5001$0.092$0.0459$509$5,092$46,424
Healthcare patient intakein7205,0003$0.310$0.0443$493$4,929$44,791
Insurance claims intake (FNOL)in9246,0004$0.388$0.0431$481$4,808$43,576
Insurance quotingin8226,0004$0.364$0.0456$506$5,055$46,053
Banking (balance, card)in384,0003$0.137$0.0458$508$5,081$46,306
Retail returns / exchangesin4113,5003$0.182$0.0456$506$5,061$46,108
Real estate lead follow-upout3.5103,0001$0.186$0.0531$581$5,811$53,609
Recruiting phone screenout10244,0002$0.486$0.0486$536$5,357$49,068
Surveys / NPSout2.592,0001$0.134$0.0537$587$5,867$54,166
Roadside assistance intakein5144,5003$0.222$0.0444$494$4,936$44,865
Travel booking changesin7186,0004$0.319$0.0456$506$5,056$46,064
Shift confirmation (internal)out131,2001$0.058$0.0578$628$6,283$58,332
Field-service dispatch (internal)in383,0003$0.137$0.0457$507$5,073$46,233

Three patterns stand out, and none of them show up in a single per-minute quote.

Per-minute cost barely changes by use case. On a lean stack, almost every meter is per-minute, so inbound use cases land between $0.043 and $0.047 per minute. The use case changes the cost per call, which scales with call length, and the cost per resolved call. It doesn't change the per-minute rate.

Outbound costs about 20–30% more per minute. Twilio's termination rate ($0.0100) is about three times its origination rate ($0.0034). Outbound calls also tend to be short, so the 60-second rounding takes a bigger bite.

Short calls cost the most per minute. Shift confirmations and reminders run $0.058 per minute, against $0.045 for support. Rounding adds about 0.5 billed minutes to a 1-minute call. That's 50% extra carrier time, compared with 11% on a 4.5-minute call.

So how much does an AI receptionist cost? On this model, about $0.09 per call. An SMB receptionist handling 5,000 two-minute calls (10k minutes) costs about $510 a month in runtime, before your engineering time. Appointment scheduling AI costs about $0.14 per call. A customer support voice agent costs about $0.21 per call, or roughly $5,000 a month at 100k minutes.

Stacked bar chart of cost per call by component for seven voice agent use cases, showing hosting and TTS as the largest shares and a cached small LLM near one percent

What moves cost the most

The chart shows the baseline split for a support call: hosting 44%, TTS 24%, telephony 17%, STT 14% and LLM 1%. That 1% is the most misleading number in this post. It only holds when you use a small model and caching works. Each meter has its own mechanics.

LLM tokens: your prompt costs more than the conversation

A cascaded agent sends the full context on every request. With S system tokens, R requests per call and h new tokens added per request, total input tokens are roughly R × S + h × R(R−1)/2. The prompt term grows linearly and the history term grows quadratically.

Plug in the support call: S = 4,000, R = 14 (12 turns plus 2 tool calls) and h ≈ 87. The prompt term is 56,000 tokens. The history term is about 7,900. The system prompt is about 88% of input tokens. For calls under 10 minutes, prompt length matters far more than conversation length.

Two consequences follow:

  • Every tool call is a hidden full-context request. Two tool calls take 12 requests to 14, which adds about 17% to input tokens. A tool that fails and gets retried costs a full round trip each time. Repetition loops have the same effect (voice agent repetition loops).
  • Trimming a 6,000-token prompt to 3,000 does more for the LLM bill than shaving a minute off the average call. See prompt engineering for voice agents for ways to compress prompts without losing behavior.

Caching: the 10x lever that breaks without warning

With caching, the static prefix bills at 0.1x and only the new tail of each request bills at the 1.25x write rate. Without caching, every request pays full input price on the whole context. The chart below shows the same support call on GPT-5.6 Luna, Terra and Sol.

Bar chart of one 4.5-minute support call across six LLM setups: Luna, Terra and Sol, each with caching working or broken; cost rises from $0.205 to $0.541

On Luna, broken caching barely matters ($0.205 to $0.216). On Terra, it takes the call from $0.224 to $0.338, a 51% jump, and the LLM grows from 9% to 40% of the call. On Sol with caching broken, the LLM is 62% of a $0.54 call.

Caching usually breaks for one of three reasons:

  • Dynamic content at the top of the prompt. If you put the current time, the caller's name or account data in the first lines of the system prompt, every call gets a new prefix. Cache reuse needs the whole rendered prefix to match. Keep instructions and tool schemas first, and put per-call data after them.
  • Prompts under the minimum. GPT-5.6 won't cache a prefix shorter than 1,024 visible tokens. A tight 900-token prompt never caches, though at that size it barely matters.
  • Cache routing at high request rates. OpenAI notes that cached states live on individual machines, and traffic above 15 requests per minute can overflow to other machines. Your hit rate can drop as you scale. Measure `input_cached_tokens` instead of assuming a rate.

The research on model choice points the same way. FrugalGPT matched GPT-4's accuracy at up to 98% lower cost by cascading cheaper models first (Chen, Zaharia and Zou, 2023). RouteLLM cut costs by more than 2x on standard benchmarks without losing response quality, by routing each query to a strong or a weak model (Ong et al., 2024). In voice, a router can send identity checks and confirmations to the small model and reserve the large model for hard turns. The catch is that you have to prove the small model handles those turns, call by call.

TTS characters

TTS is the largest model cost on a cached small model: $0.049 of a $0.205 support call. It scales with what the agent says, not with call length. Three common habits inflate it: long greetings and disclosures, reading back whole confirmations ("So that's Tuesday the 14th at 3:30 pm at the Main Street location, is that right?"), and verbose fallback lines that repeat on every misunderstanding.

The vendor choice matters more than any of these. At $0.08 per 1,000 characters (ElevenLabs v3), the same support call costs $0.286 instead of $0.205. ElevenLabs lists v4 Turbo at $0.011 per 1,000 characters until October 12, 2026, down from $0.04. Model your budget at the regular rate, not the promo.

Telephony rounding and outbound attempts

Inbound SIP minutes are cheap. Rounding and failed outbound attempts are where telephony money goes. On outbound campaigns, unanswered and voicemail attempts still bill a carrier minute. If your dialer keeps an agent session open while the call rings or while voicemail detection runs, they bill session and STT time too.

Here's a worked example for appointment reminders. Assume a 35% connect rate, and assume each failed attempt bills one carrier minute plus 30 seconds of agent time. That adds about 1.86 failed attempts per connected call. Cost per connected reminder rises from $0.070 to about $0.117, and cost per confirmed appointment from $0.082 to about $0.137. Start the agent session only after the call is answered, and keep voicemail detection fast (voicemail detection on LiveKit and Pipecat).

Hosting and observability

On LiveKit Cloud, the agent session ($0.01/min) and observability ($0.01/min) are 44% of a support call. Observability is the single largest line you can turn off. Dropping it takes the support call from $0.205 to $0.160, a 22% cut. Don't drop it blind, though. If you sample sessions instead of recording all of them, keep enough recorded calls to catch regressions. Telephony and hosting choices move this line more than model choices do.

Sensitivity: change one component at a time

Each row changes one component and holds everything else at baseline. The support call is 4.5 minutes inbound, the tech support call is 8 minutes inbound, and the reminder is 1.2 minutes outbound.

ScenarioSupport callSupport per minTech support callReminder call
Baseline (Flux, Luna cached, Aura-2, Twilio SIP, LiveKit Cloud)$0.205$0.0455$0.364$0.070
LLM: GPT-5.6 Luna, caching broken$0.216$0.0481$0.395$0.071
LLM: GPT-5.6 Terra, cached$0.224$0.0497$0.407$0.074
LLM: GPT-5.6 Terra, caching broken$0.338$0.0751$0.711$0.089
LLM: GPT-5.6 Sol, cached$0.255$0.0566$0.480$0.081
LLM: GPT-5.6 Sol, caching broken$0.541$0.1201$1.240$0.118
TTS: ElevenLabs Flash/Turbo ($0.04/1k chars)$0.221$0.0491$0.393$0.075
TTS: Cartesia Sonic at Scale overage ($0.038/1k)$0.218$0.0484$0.387$0.074
TTS: ElevenLabs v3 ($0.08/1k)$0.286$0.0635$0.508$0.096
STT: Deepgram Nova-3 instead of Flux$0.197$0.0438$0.351$0.068
Telephony: Telnyx SIP ($0.0032 in, $0.005 out)$0.204$0.0453$0.362$0.061
Drop LiveKit observability$0.160$0.0355$0.284$0.058
Pipecat + Daily PSTN ($0.025/min), self-hosted at $0.002/min (assumed)$0.214$0.0475$0.372$0.069
Volume rates: Flux and Aura-2 Growth, SIP $0.003, Telnyx$0.191$0.0424$0.339$0.058

The Cartesia rate is the Scale plan overage of $38 per 1M credits, at 1 credit per character (Cartesia pricing). The Daily row uses $0.025 per PSTN minute plus $2 a month per US number, and an assumed $0.002 per minute for your own compute. Telnyx rates are from its SIP pricing page.

Two takeaways:

  • The LLM decision can triple the bill on long calls. On tech support, Sol without caching costs 3.4x the baseline. Nothing else in the table moves cost that much.
  • Telephony vendors barely matter on inbound but matter on outbound. Telnyx saves $0.001 on a support call and 13% on a reminder.

One caveat: Nova-3 is cheaper than Flux, but Flux includes end-of-turn detection. Swapping STT changes turn-taking, which changes call length and resolution. A model swap is never only a price change. Run the same scenarios through both configurations before you switch. STT cost vs accuracy covers that trade-off.

Cost per resolved call

Cost per call tells you what the platform bill will be. Cost per resolved call tells you whether the agent pays for itself. The formula is cost per call ÷ resolution rate. The resolution rates below are illustrative assumptions, not benchmarks. Replace them with your own measured rate (how to measure resolution rate).

Use caseAssumed resolutionCost/resolved call
Customer support (general)55%$0.373
Technical support / troubleshooting40%$0.910
Appointment scheduling75%$0.183
Appointment reminders85%$0.082
Order status (WISMO)80%$0.118
Food ordering70%$0.255
Restaurant reservations80%$0.115
Billing inquiries55%$0.332
Collections (early-stage)30%$0.635
Inbound lead qualification60%$0.295
Outbound sales / appt setting15%$1.117
SMB front-desk receptionist70%$0.131
Healthcare patient intake65%$0.477
Insurance claims intake (FNOL)55%$0.705
Insurance quoting35%$1.041
Banking (balance, card)65%$0.211
Retail returns / exchanges65%$0.281
Real estate lead follow-up50%$0.372
Recruiting phone screen70%$0.694
Surveys / NPS60%$0.224
Roadside assistance intake75%$0.296
Travel booking changes45%$0.709
Shift confirmation (internal)90%$0.064
Field-service dispatch (internal)80%$0.171
Paired bar chart comparing cost per call and cost per resolved call for eight voice agent use cases, where low resolution rates turn cheap outbound sales calls into the most expensive outcomes

The order flips. Outbound sales is one of the cheaper calls at $0.167, but one of the most expensive outcomes at $1.12. Resolution is the most powerful lever in the model. For the support call, moving resolution from 40% to 70% cuts cost per resolved call from $0.512 to $0.293. No vendor swap in the sensitivity table comes close.

These figures still leave out the human who handles the calls the agent doesn't resolve. Add that cost to get true cost per outcome. The full method is in voice agent cost per resolution, and the payback math is in voice agent ROI.

Monthly cost and capacity at 10k, 100k and 1M minutes

The monthly columns above are minutes × per-minute cost + plan fee. Capacity is the part people forget. Your plan has to support peak concurrency, not average concurrency.

Here's a worked example, assuming traffic over 22 business days of 10 hours each, which is 13,200 open minutes a month:

  • 10k minutes/month: average concurrency is 10,000 ÷ 13,200 ≈ 0.8. With a 2.5x peak-hour factor (assumed), the peak is about 2 calls. LiveKit's Build plan allows 5 concurrent sessions and Ship allows 20.
  • 100k minutes/month: average concurrency is about 7.6, and the peak is about 19. That's right at Ship's 20-session limit, so plan for Scale ($500 a month, up to 600 sessions).
  • 1M minutes/month: average concurrency is about 76, and the peak is about 190. That's Scale, or Enterprise if you need more. Check your STT and TTS concurrency limits too. Deepgram's pay-as-you-go tier allows up to 150 streaming STT connections and 45 TTS connections (Deepgram pricing). At 1M minutes, you'll hit the TTS limit before the LLM bill gets interesting.

Fixed monthly costs are small next to usage. Phone numbers run $1.00 to $1.15 a month on Telnyx and Twilio, and the LiveKit plan fee is $50 or $500. At 1M minutes, volume discounts are worth chasing. The "volume rates" row in the sensitivity table saves about 7% per call. The bigger savings at that scale come from how many tokens and characters each call uses, not from discounts. Build vs buy voice agent cost covers when that scale justifies running your own stack.

A Python voice AI cost calculator

This is the model that produced every number above. It's illustrative, so replace the prices with your contract rates and the use case fields with your own call data.

"""voice_cost.py - illustrative per-call cost model for an in-house voice agent.
Prices: public list rates checked Oct 5, 2026. Every workload number is an assumption."""
import math
from dataclasses import dataclass

P = dict(
    stt_min=0.0065,          # Deepgram Flux English, streaming, per audio minute
    tts_1k_chars=0.030,      # Deepgram Aura-2, per 1,000 characters
    llm_in=0.20e-6,          # GPT-5.6 Luna, uncached input, per token
    llm_cache_read=0.02e-6,  # cached input (0.1x)
    llm_cache_write=0.25e-6, # cache write (1.25x uncached) on GPT-5.6
    llm_out=1.20e-6,         # output, per token
    session_min=0.01,        # LiveKit Cloud agent session minute
    observ_min=0.01,         # LiveKit agent observability minute
    sip_min=0.004,           # LiveKit third-party SIP minute (Ship plan)
    carrier_in=0.0034,       # Twilio Elastic SIP origination, US local
    carrier_out=0.0100,      # Twilio Elastic SIP termination, lower 48
)
CHARS_PER_SPOKEN_MIN = 800   # Cartesia pricing FAQ: 750-800 chars per audio minute
CHARS_PER_TOKEN = 4
USER_TOKENS = 30             # tokens per caller turn (assumption)
TOOL_TOKENS = 225            # tool-call JSON + result per tool call (assumption)

@dataclass
class UseCase:
    name: str
    direction: str     # "in" or "out"
    minutes: float     # average connected call length
    turns: int         # agent turns per call
    sys_tokens: int    # system prompt + tool schemas
    tools: int         # tool calls per call (each adds one LLM request)
    talk: float        # share of call time the agent is speaking
    resolution: float  # share of calls resolved without a human

def llm_cost(u, p=P, cache=True):
    out_per_turn = u.talk * u.minutes * CHARS_PER_SPOKEN_MIN / CHARS_PER_TOKEN / u.turns
    requests = u.turns + u.tools
    new_per_req = (u.turns * (USER_TOKENS + out_per_turn) + u.tools * TOOL_TOKENS) / requests
    cost, history = 0.0, 0.0
    for _ in range(requests):
        prefix = u.sys_tokens + history
        if cache:   # prefix read from cache, new tail written to cache
            cost += prefix * p["llm_cache_read"] + new_per_req * p["llm_cache_write"]
        else:
            cost += (prefix + new_per_req) * p["llm_in"]
        history += new_per_req
    return cost + u.turns * out_per_turn * p["llm_out"]

def call_cost(u, p=P, cache=True, round_carrier=True):
    m = u.minutes
    billed = m + 0.5 if round_carrier else m  # 60/60 rounding adds ~0.5 min per call on average
    carrier = p["carrier_in"] if u.direction == "in" else p["carrier_out"]
    parts = {
        "stt": m * p["stt_min"],
        "tts": u.talk * m * CHARS_PER_SPOKEN_MIN / 1000 * p["tts_1k_chars"],
        "llm": llm_cost(u, p, cache),
        "hosting": m * (p["session_min"] + p["observ_min"]),
        "telephony": billed * carrier + m * p["sip_min"],
    }
    parts["total"] = sum(parts.values())
    return parts

def monthly_cost(u, minutes_per_month, p=P, fixed=0.0):
    return minutes_per_month / u.minutes * call_cost(u, p)["total"] + fixed

def outbound_cost_per_contact(u, p=P, connect_rate=0.35, failed_attempt_min=0.5):
    """Unanswered and voicemail attempts still bill a full carrier minute, plus agent time."""
    failed = (1 - connect_rate) / connect_rate
    per_failed = (math.ceil(failed_attempt_min) * p["carrier_out"]
                  + failed_attempt_min * (p["stt_min"] + p["session_min"]
                                          + p["observ_min"] + p["sip_min"]))
    return call_cost(u, p)["total"] + failed * per_failed

if __name__ == "__main__":
    support = UseCase("Customer support", "in", 4.5, 12, 4000, 2, 0.45, 0.55)
    c = call_cost(support)
    print({k: round(v, 4) for k, v in c.items()})
    print("per resolved:", round(c["total"] / support.resolution, 3))
    print("100k min/month:", round(monthly_cost(support, 100_000, fixed=500)))

Running it prints the baseline support call: STT $0.0292, TTS $0.0486, LLM $0.0021, hosting $0.09 and telephony $0.035, for a total of $0.2049. Per resolved call it's $0.373, and at 100k minutes it's $5,054 a month. To model Terra, set `llm_in=2e-6`, `llm_cache_read=0.2e-6`, `llm_cache_write=2.5e-6` and `llm_out=12e-6`. To see what broken caching costs, pass `cache=False`.

How to measure your real cost per call

A model is a hypothesis. These steps replace each assumption with a measurement from your own traffic.

1. Log per-session usage from the SDK. In LiveKit Agents, `session.usage.model_usage` holds cumulative per-model totals. LLM entries carry `input_tokens`, `input_cached_tokens` and `output_tokens`. TTS entries carry `characters_count` and `audio_duration`. STT entries carry `audio_duration` (LiveKit data hooks). The older session-level `metrics_collected` event and `UsageCollector` are deprecated. A simplified handler:

import json, logging
from livekit.agents import AgentServer, AgentSession, JobContext

logger = logging.getLogger("cost")
server = AgentServer()

@server.rtc_session(agent_name="support-agent")
async def entrypoint(ctx: JobContext):
    session = AgentSession()  # configure stt/llm/tts as usual (simplified)

    async def log_usage():
        row = {"room": ctx.room.name, "models": []}
        for u in session.usage.model_usage:
            entry = {"provider": u.provider, "model": u.model}
            if hasattr(u, "characters_count"):          # TTSModelUsage
                entry["tts_chars"] = u.characters_count
            elif hasattr(u, "input_cached_tokens"):     # LLMModelUsage
                entry["in"] = u.input_tokens
                entry["cached"] = u.input_cached_tokens
                entry["out"] = u.output_tokens
            elif hasattr(u, "audio_duration"):          # STTModelUsage
                entry["stt_seconds"] = u.audio_duration
            row["models"].append(entry)
        logger.info(json.dumps(row))

    ctx.add_shutdown_callback(log_usage)
    await ctx.connect()
    # await session.start(...)

2. Compute your cache hit rate per call as `cached ÷ in`. A rate far below (S × R) ÷ total input means something dynamic sits in your prefix. Uncached tokens on GPT-5.6 include cache writes at 1.25x, so price uncached input at the write rate for an upper bound.

3. Pull carrier minutes from your Twilio or Telnyx usage records, not from session length. Compare billed minutes with connected minutes per call to get your real rounding overhead. For outbound, count attempts per connected call.

4. Replace the model's workload fields with medians and p90s from about 500 recent calls per use case: minutes, agent turns, tool calls, TTS characters. Costs are skewed, so check the p90 as well as the median. A few long calls can drive the monthly bill.

5. Measure resolution on a scored sample, not a dashboard flag. A call marked "completed" isn't a resolved call. Score a sample against a written definition of resolved for each use case (containment vs resolution). This is where an independent evaluation like Evalgent's fits: it scores resolution the same way across configurations, so your cost per resolved call rests on a number nobody tuned to look good.

6. Re-run cost and resolution together after every prompt or model change. A shorter prompt that saves $0.01 per call but drops resolution by five points costs you money. Gate changes on both numbers (ship prompt changes safely). Latency is the third axis, because slower agents run longer calls and lose callers (cost of latency).

Frequently asked questions

What is the AI voice agent cost per call?

On an in-house stack at October 2026 list prices, expect about $0.045 per inbound minute and $0.053–$0.058 per outbound minute. That puts a 2-minute receptionist call at about $0.09, a 4.5-minute support call at about $0.21 and a 9-minute claims intake at about $0.39. Your call length matters more than your use case label.

How much does an AI receptionist cost per month?

At 10,000 minutes a month, which is about 5,000 two-minute calls, runtime costs about $510 on the baseline stack, including a $50 LiveKit plan. At 100,000 minutes it's about $5,100. Engineering time, monitoring and human escalation aren't included. At low volume, those can cost more than the runtime.

What is the biggest cost driver in a voice agent call?

On a cached small LLM, it's hosting plus observability (about 44%), then TTS (about 24%). On a frontier LLM with broken prompt caching, the LLM takes over at 40–70% of the call. The system prompt is resent on every request, so it accounts for most input tokens on calls under 10 minutes.

How do I calculate cost per resolved call?

Divide cost per call by your measured resolution rate. A $0.205 support call at 55% resolution costs $0.373 per resolved call. At 40% it costs $0.512, and at 70% it costs $0.293. Add the human cost of the unresolved calls if you want true cost per outcome.

Does prompt caching really matter for voice agents?

It depends on the model. On GPT-5.6 Luna, broken caching adds about 5% to a support call. On GPT-5.6 Terra, it adds 51%. On Sol, it more than doubles the call. Keep static instructions and tool schemas at the top of the prompt, and put per-call data after them.

Why do short calls cost more per minute?

Twilio and Telnyx round each call up to a full minute. A 61-second reminder is billed as 2 minutes. On a 1-minute call, rounding adds about 50% to carrier time, compared with about 11% on a 4.5-minute call. Failed outbound attempts add carrier minutes that produce no conversation at all.

Is Pipecat with Daily cheaper than LiveKit Cloud?

On this model, the two are close. Daily PSTN at $0.025 a minute replaces LiveKit's session, observability, SIP and carrier lines, and the support call comes to $0.214 vs $0.205. The result depends on what your own hosting and observability really cost, so price both with your traffic.

How accurate are these voice agent cost estimates?

The prices are verified list rates. The workloads are assumptions. Treat these figures as starting estimates until you replace call length, turns, prompt size and resolution with your own medians. The measurement steps above show how to do that from SDK usage data.

The bottom line

A voice agent's per-minute cost barely changes across use cases, so budget per call and judge per resolved call. Measure your real token, character and carrier usage, then spend your effort on prompt size, caching and resolution rate, because those move the number more than any vendor swap.

Related Articles