Evalgent
Back to Blog
Voice AI Evaluation

Best LLM for voice agents (October 2026): latency, tool calling, cost per minute

Deepesh Jayal
Updated
17 min read
Best LLM for voice agents (October 2026): latency, tool calling, cost per minute
On this page

Prices and model names in this guide were checked against official pricing pages on October 5, 2026. Models change monthly, so recheck before you commit.

Most "best LLM for voice" lists rank models by intelligence scores. That is the wrong axis. A voice agent spends its life in short, many-turn, tool-heavy conversations where a 900 ms pause sounds like a broken line. The model that tops a reasoning leaderboard can be the worst choice for a scheduling bot.

This guide is for teams that build their own pipeline, usually on LiveKit, Pipecat or Dograh, with STT and TTS wired in by hand. It covers what to measure, how to read published benchmarks without being misled, what each model family costs per minute once caching is on, and how to run a bake-off on your own calls. If you are still deciding on the rest of the stack, start with the voice agent stack guide.

The October 2026 shortlist at a glance

These are the text models worth testing for a cascaded STT, LLM, TTS pipeline, with list prices per million tokens from each vendor's pricing page.

ModelInput / output per 1MCached inputVoice-relevant note
GPT-6 Luna (`gpt-6-luna`)$0.10 / $0.50$0.01Supports `reasoning.effort: "none"`
GPT-6.1 Sol (`gpt-6.1-sol`)$2.00 / $10.00$0.10Lowest effort is `low`; `none` not supported
GPT-6 Astra (`gpt-6-astra`)$10.00 / $50.00$1.00Ultrafast tier costs 6x the standard price
GPT-5.6 Sol$4.00 / $20.00 (promo)$0.40Still offered; promo pricing through at least Nov 21, 2026
Claude Haiku 4.5$1 / $50.1x inputAnthropic's fastest model
Claude Sonnet 5.5$2 / $100.1x input"Best combination of speed and intelligence"
Claude Opus 5.5$4 / $200.05x inputFlagship; overkill for most live turns
Gemini 3.8 Flash$0.75 / $3.75$0.075Intro price; doubles on Jan 1, 2027
Gemini 3.5 Flash-Lite$0.30 / $2.50$0.03Cheapest Google option
Grok 4.7$2.00 / $6.00$0.50Configurable reasoning
GPT-OSS-120B on Groq$0.15 / $0.60n/aGroq lists about 500 tokens/sec
GPT-OSS-20B on Groq$0.075 / $0.30n/aGroq lists about 1,000 tokens/sec

Sources: OpenAI pricing, Anthropic pricing, Gemini API pricing, xAI models, GroqCloud models.

Two details in that table matter more than the prices. First, OpenAI's GPT-6 guide says GPT-6 Astra and GPT-6.1 Sol do not support the `none` reasoning effort, while GPT-6 Luna and the older GPT-6 Sol do. So the cheapest GPT-6 model is also the only current one you can run with no thinking at all. Second, Gemini's pricing page bills output "including thinking tokens," so a Flash model left on its default thinking setting costs more and starts talking later than the list price suggests.

Why time to first token beats throughput for voice

A voice turn does not wait for the full LLM reply. Your framework streams tokens into the TTS engine, and the TTS can start speaking once it has a speakable chunk, usually the first clause or sentence. So the latency the caller hears from the LLM is:

LLM share of turn latency = TTFT + (tokens in first speakable chunk ÷ output speed)

TTFT is made of network round trip, queueing, prefill (processing your input tokens), and any reasoning tokens generated before the answer. Output speed only matters for the 10 to 20 tokens of the first chunk. After that, the TTS plays audio far slower than any modern model generates text, so the rest of the reply streams in while the caller is still listening.

Worked example, illustrative numbers. Take a first sentence of 15 tokens:

  • Model A: TTFT 350 ms, 80 tokens/sec. First chunk ready at 350 + (15 ÷ 80 × 1000) = 350 + 188 = 538 ms.
  • Model B: TTFT 700 ms, 300 tokens/sec. First chunk ready at 700 + 50 = 750 ms.

Model B is almost four times faster on throughput and still loses by 212 ms. Past roughly 100 tokens/sec, extra throughput barely moves voice latency. TTFT does.

Now add reasoning. If a model emits 200 hidden reasoning tokens at 150 tokens/sec before the answer, that is 200 ÷ 150 = 1.33 seconds of added silence on every turn. That is why "reasoning off on the live turn" is the most important setting in this whole guide. Humans leave gaps of roughly 200 ms between turns across languages, per Stivers et al. in PNAS. Your whole pipeline, not just the LLM, has to land close enough to that for the call to feel normal. The cost of latency post shows what slow turns do to abandonment.

Stacked latency bars comparing two LLMs for a voice turn, showing network, prefill, reasoning and first-sentence decode, with the high-throughput model finishing its first sentence later

Tool turns pay TTFT twice

When the model calls a tool, the turn needs two LLM requests. The first emits the function call; your code runs the tool; the second turns the result into speech. Turn latency becomes:

Tool turn = TTFT₁ + tool-call tokens ÷ speed + tool execution + TTFT₂ + first chunk ÷ speed

With a 400 ms TTFT and a 300 ms CRM lookup, you are past 1.1 seconds before the caller hears anything. That is why a slow-TTFT model can test fine on FAQ calls and feel broken on booking calls. Test both kinds of turn separately, and measure tool turns as their own latency bucket. Our time to first audio guide explains how to instrument this end to end.

Reading published latency numbers without being misled

Artificial Analysis is the most widely used independent source for LLM API latency. Its methodology has four details that change how a voice team should read it.

1. The default workload is a 10,000-token prompt. Since March 2026 the site's default speed numbers come from 10k-input prompts (previously 1k). A voice turn with a cached 3,000-token system prompt has a much smaller uncached prefill. Switch the prompt option to 1k to get closer to voice conditions.

2. Numbers are medians, not tails. Results are the P50 over the past 72 hours. Callers remember the slow turns. You need your own p95, measured from your region.

3. TTFT on reasoning models counts the first reasoning token. The site separates "Time to First Token" from "Time to First Answer Token." For voice, only first answer token matters, since reasoning tokens are not speakable.

4. It tests from one place. The primary test server is in Google Cloud `us-central1-a`, and network latency is part of TTFT. If your agent servers run in `us-east-1`, your numbers will differ.

The reasoning trap is easy to see on a real page. On October 5, 2026, Artificial Analysis listed GPT-6 Luna (max), the maximum-reasoning variant, with a TTFT of 95.97 seconds against a 2.17 second median for similarly priced reasoning models. Nobody would put that setting on a phone line. The same model with `reasoning.effort: "none"` is a different product for latency purposes. Always check which variant and effort level a leaderboard row describes before you compare it.

Tool calling: what the benchmarks show and what they hide

Two public benchmarks matter most for voice agents that take actions.

τ-bench (Yao et al., 2024) simulates a user talking to an agent that has tools and a policy document, in retail and airline domains. Its biggest contribution is the pass^k metric: the chance an agent succeeds on the same task in all k independent tries. In the original paper, the best GPT-4o function-calling agent had over 60% average success, yet pass^8 in retail fell below 25%. A model that solves a task "most of the time" will still fail many of your callers.

τ²-bench (Barres et al., 2025) adds a telecom domain where the user also has tools and must be guided through actions, like a support call where the caller has to toggle a phone setting. The authors found significant performance drops when agents moved from solving tasks alone to guiding a user. That is the voice support situation exactly.

BFCL, the Berkeley Function Calling Leaderboard (Patil et al., ICML 2025), checks call structure with AST matching and includes multi-turn and irrelevance categories. Its central finding is that top models do well on single-turn calls, while memory, dynamic decisions and long-horizon reasoning are still open problems.

What these hide, for voice:

  • Saturation. Third-party aggregators now list several 2025 to 2026 flagships in the high 90s on τ²-bench telecom. At that point the leaderboard no longer separates the models you are choosing between.
  • Settings. Vendor-reported scores are usually run with reasoning on at high effort. You will run with reasoning off or low. Your tool accuracy at your settings may be much lower than the headline.
  • Clean text input. Benchmarks feed perfect text. Your model gets STT output: "my zip is nine oh two one oh" or a mangled last name. Argument accuracy on noisy transcripts is a separate skill. The tool argument accuracy post covers how to score it.

Instruction following over a 20-turn call

Voice calls are long, underspecified conversations. Callers add details one at a time and change their minds. Laban et al. (2025) ran over 200,000 simulated conversations and found that every top open and closed model they tested did worse when the same task was spread across turns, with an average drop of 39% across six tasks. Most of the drop came from unreliability, not lost skill. Models made early assumptions, committed to them, and did not recover.

For a voice agent, that shows up as:

  • confirming an appointment time before the caller has given a date,
  • carrying a wrong account number forward after the caller corrected it,
  • dropping a policy rule (such as "never quote a price before verifying identity") by turn 15.

The fix is partly prompting and partly model choice, but you only see it with multi-turn tests. A single-turn eval set will rate every candidate as fine. Per-turn reliability also compounds: a model that follows a rule 99% of the time per turn keeps it for all 10 relevant turns only 0.99¹⁰ ≈ 90.4% of the time.

Prompt caching: the biggest lever on cost and a real one on latency

Every turn of a voice call resends the system prompt, tool schemas and the growing history. Without caching, you pay full price for the same 3,000 tokens a dozen times per call, and the model re-prefills them each turn. Caching reduces both.

OpenAI's prompt caching guide lists the current rules for GPT-5.6 and later:

  • Cache writes cost 1.25x the input rate. Reads cost 0.1x (0.05x on GPT-6.1 Sol).
  • The minimum cacheable prefix is 1,024 visible tokens.
  • Cached prefixes last at least 30 minutes after the last write or reuse (`prompt_cache_options.ttl: "30m"`).
  • `prompt_cache_options.prewarm: true` writes the cache without generating output, specifically to cut TTFT on the next request.

The voice-specific gotchas follow from "the whole rendered prefix must match":

1. Never put the caller's name, the current time or the account ID at the top of the system prompt. It changes per call and breaks the shared prefix for every caller. Put stable instructions first and dynamic context in a later message.

2. Do not add or remove tools mid-call. Changing tool names, schemas or order changes the prefix. OpenAI recommends keeping the `tools` list fixed and using `allowed_tools` or `tool_choice: "none"` to restrict calls.

3. Do not change `reasoning.effort` at the request level mid-call. On GPT-6 models, append a `configuration_update` input item instead, which keeps the cached prefix.

4. Prewarm before you answer. Fire a prewarm request when the inbound call event arrives, so the cache is warm by the time the caller finishes saying hello.

5. Watch compaction and history trimming. Summarizing old turns rewrites the prefix and resets reuse for the next request.

Anthropic and Google offer similar discounts: Anthropic charges 1.25x for 5-minute cache writes and about 0.1x for reads (0.05x on Opus 5.5), and Gemini 3.8 Flash charges $0.075 per million cached tokens plus an hourly storage fee. Each vendor has its own breakpoint and minimum-length rules, so verify cache hits in the usage fields rather than assuming them.

Cost per minute, worked out

Assumptions for this worked example (replace with your own): a 5-minute call, 12 LLM turns, a 3,000-token system prompt plus tool schemas, history growing 150 tokens per turn, 60 output tokens per turn. That is 45,900 input tokens and 720 output tokens per call. With caching, about 41,250 input tokens are cache reads and 4,650 are cache writes.

ModelUncached, per minuteCached, per minute
GPT-6 Luna$0.0010$0.0003
GPT-OSS-120B on Groq$0.0015n/a
Gemini 3.8 Flash$0.0074$0.0019
Claude Haiku 4.5$0.0099$0.0027
GPT-6.1 Sol$0.0198$0.0046
Claude Sonnet 5.5$0.0198$0.0054
Grok 4.7$0.0192$0.0068
Claude Opus 5.5$0.0396$0.0092
GPT-6 Astra$0.0990$0.0271
GPT-6 Astra, Ultrafast$0.5940$0.1625

Formula per call: (uncached input × input price + cache reads × read price + cache writes × write price + output × output price) ÷ 1,000,000. Gemini storage fees and long-context surcharges are excluded.

Three things fall out of this table:

  • Caching cuts LLM cost by roughly 65 to 77% on every model that supports it, because voice prompts are dominated by repeated input.
  • For most models the LLM is now a small slice of the per-minute bill. Below a cent a minute, STT, TTS and telephony usually cost more. Use the voice agent pricing comparison for the other lines.
  • The exception is the top tier. Astra on Ultrafast at about $0.16 per cached minute is more than three times what GPT-Live charges for a whole speech-to-speech session. Reserve the flagship for the few turns that need it.
Bar chart of worked-example LLM cost per call minute for ten models, uncached versus cached, showing caching cutting cost by two thirds to three quarters

Speech-to-speech options in October 2026

If you skip the cascade, the "LLM" choice becomes a realtime model choice. Current options with official pricing:

OptionPriceNote
OpenAI GPT-Live (`gpt-live-1`)$0.05/min, billed per secondFull-duplex; delegates reasoning and tools to a backend Responses model billed separately
OpenAI `gpt-realtime-2.1`$32 / $64 per 1M audio tokensToken-billed realtime model
OpenAI `gpt-realtime-2.1-mini`$10 / $20 per 1M audio tokensCheaper realtime tier
Gemini 3.8 Liveabout $0.005/min audio in, $0.018/min audio outOptional Extended Thinking variant
xAI Grok Voice$0.08/min+$0.01/min telephony on an xAI number

Two gotchas. GPT-Live's $0.05 is not the full price: per the model page, backend model and tool usage is billed separately, so a Luna backend adds little but a Sol or Astra backend adds the cached rates above. And GPT-Live rate limits are counted in concurrent sessions, starting at 25 on Tier 1, which caps an outbound campaign before cost does. The OpenAI Realtime and Gemini Live phone calls guide covers the telephony side, and evaluating a realtime voice API covers testing.

Selection matrix by use case

Use caseStart withEscalate toWhat to test hardest
FAQ, hours, order statusGPT-6 Luna, Gemini 3.5 Flash-Lite, GPT-OSS-20BRarely neededp95 TTFT, refusal to invent answers
Scheduling, reminders, shift confirmationGPT-6 Luna, Claude Haiku 4.5, Gemini 3.8 FlashGPT-6.1 Sol on reschedule conflictsDate and time argument accuracy, multi-turn corrections
Order taking, lead qualificationClaude Haiku 4.5, Gemini 3.8 Flash, GPT-OSS-120BSonnet 5.5 or GPT-6.1 SolSlot filling across 10+ turns, upsell policy
Complex support, troubleshootingGPT-6.1 Sol (low), Claude Sonnet 5.5, Grok 4.7Opus 5.5 or Astra, offline or asyncτ²-style guided actions, recovery after wrong turn
Regulated: collections, healthcare intake, insuranceSonnet 5.5 or GPT-6.1 Sol, with a small classifier for policy checksHuman handoff, not a bigger modelDisclosure wording, identity verification order, pass^k on every rule

For regulated flows, the model is not your compliance control. Script-level rules need deterministic checks around the model and test cases for every rule, as in regulated voice agent policy test cases.

Decision matrix mapping five voice agent use cases to a starting model tier, an escalation tier, and the test that matters most for each

Routing: a small model by default, a big model by exception

You do not have to pick one model. The pattern most teams settle on is:

1. A fast, no-reasoning model handles every turn by default.

2. A router classifies each turn (or each intent) and sends hard turns, such as billing disputes, multi-constraint rescheduling, or anything after two failed attempts, to a stronger model.

3. The voice keeps flowing while the big model works: speak a short filler from the small model or a fixed phrase, then hand the bigger model's answer to TTS.

On GPT-6 you can also raise reasoning effort for a single hard turn with a `configuration_update` item and lower it again, without losing the cache. The router itself has to be fast and highly consistent, or it adds the latency you were trying to save. The model routing for voice agents post goes deeper on router design, including using a classifier such as Jev, a highly consistent model at $0.042 per million input tokens with 70 to 500 ms responses.

Measuring TTFT and first sentence yourself

Simplified Python against the OpenAI Responses API. It streams a reply with no reasoning and records time to first answer token and time to first sentence. Run it from the same region as your agent servers, at least 200 times per model, and keep p50 and p95.

# Simplified: measure TTFT and time-to-first-sentence for one turn.
import time, statistics
from openai import OpenAI

client = OpenAI()
SYSTEM = open("system_prompt.txt").read()   # stable prefix, 1,024+ tokens

def one_turn(user_text: str, model: str = "gpt-6-luna"):
    t0 = time.perf_counter()
    ttft = first_sentence = None
    buf = ""
    stream = client.responses.create(
        model=model,
        reasoning={"effort": "none"},          # Luna supports "none"
        input=[
            {"role": "developer", "content": SYSTEM},
            {"role": "user", "content": user_text},
        ],
        stream=True,
    )
    for event in stream:
        if event.type == "response.output_text.delta":
            now = time.perf_counter()
            if ttft is None:
                ttft = now - t0
            buf += event.delta
            if first_sentence is None and any(p in buf for p in ".?!"):
                first_sentence = now - t0
        elif event.type == "response.completed":
            u = event.response.usage
            cached = u.input_tokens_details.cached_tokens
    return ttft, first_sentence, cached

runs = [one_turn("Can I move my Tuesday appointment to Thursday?") for _ in range(200)]
ttfts = sorted(r[0] for r in runs)
print("p50", statistics.median(ttfts), "p95", ttfts[int(0.95 * len(ttfts)) - 1])
print("cache hit rate", sum(r[2] > 0 for r in runs) / len(runs))

Check the cached-token count on every run. If the hit rate is low, your prefix is changing between requests, and your TTFT numbers are measuring the uncached case.

How to evaluate an LLM for your own voice agent

1. Build a test set from real calls. Pull 100 to 300 transcripts and turn them into scenarios: about 40% routine paths, 30% tool-heavy paths, 20% corrections and mind changes, 10% adversarial or out-of-scope. A golden dataset should include the STT errors you actually see.

2. Freeze everything except the model. Same prompt, same tools, same STT and TTS, same region. Only the LLM and its reasoning setting change.

3. Define the metrics before you run. Task success (did the booking exist in the system at the end), tool-call accuracy (right function, right arguments), policy adherence per rule, p50 and p95 time to first audio, and cost per resolved call.

4. Run each scenario k times. Use k = 4 to 8 and report pass^k, not just average success. A model at 92% average and 70% pass^4 is less dependable than one at 89% and 82%.

5. Size the comparison honestly. To tell 90% from 95% task success at 80% power and α = 0.05, you need about 435 scenario runs per model: n ≈ [1.96√(2 × 0.925 × 0.075) + 0.84√(0.9 × 0.1 + 0.95 × 0.05)]² ÷ 0.05². Smaller differences need more.

6. Listen to the disagreements. Where models differ on the same scenario, play both calls. Text graders miss pacing, interruptions and tone, which is why an LLM judge alone has limits.

7. Re-run on every model update. Vendors ship new snapshots monthly. Pin versions where you can, and treat each upgrade as a release, as in LLM update regressions.

This is the part most teams skip, because hand-dialing 400 calls per model is not realistic. It is also where independent evaluation helps most. Evalgent runs the same simulated callers against each candidate, with accent and noise profiles, and scores latency, tool accuracy and policy adherence with thresholds you set. You can keep using the same suite as a regression gate after you choose. The broader method is in the AI voice agent testing guide.

Frequently asked questions

What is the best LLM for voice agents in October 2026?

There is no single winner. For most cascaded voice agents, start with a fast model with reasoning off: GPT-6 Luna, Claude Haiku 4.5, Gemini 3.8 Flash or GPT-OSS-120B on Groq. Escalate hard turns to GPT-6.1 Sol, Claude Sonnet 5.5 or Grok 4.7. Pick the final model by testing on your own scenarios.

Why does TTFT matter more than tokens per second?

TTS starts speaking once the first clause arrives, so the caller only waits for TTFT plus 10 to 20 tokens. At 100 tokens per second or more, those tokens take 100 to 200 ms. A slow first token adds its full delay to every turn, while extra throughput after that is mostly hidden behind playback.

Should reasoning be turned on for voice agents?

Not on the live turn by default. Reasoning tokens arrive before any speakable text, so 200 reasoning tokens at 150 tokens per second add over a second of silence. Use reasoning for offline work or for single hard turns, raised temporarily with a configuration update on GPT-6 so the cache survives.

How much does the LLM cost per minute of a voice call?

In our worked example of 12 turns over 5 minutes with a 3,000-token prompt, cached LLM cost ranges from about $0.0003 per minute on GPT-6 Luna to $0.027 on GPT-6 Astra. Most mid-tier models land between $0.002 and $0.007. Caching cuts these numbers by two thirds to three quarters.

How does prompt caching help voice agent latency?

Cached tokens skip prefill, so the model starts generating sooner, and they cost 0.05x to 0.1x the input price. To get hits, keep the system prompt and tool list identical across calls, move caller-specific data later in the prompt, and prewarm the cache while the phone is ringing.

Is GPT-Live cheaper than a cascaded pipeline?

GPT-Live costs $0.05 per minute, billed per second, plus separate charges for its backend model and tools. A cascaded pipeline with a small LLM can cost less on the LLM line, but you also pay for STT and TTS. Compare full per-minute cost and test both on the same scenarios.

Which benchmark best predicts tool calling on calls?

τ²-bench is the closest public proxy because it simulates users who must be guided through actions, and τ-bench's pass^k metric captures consistency. Both use clean text and usually high reasoning effort, so they overstate accuracy at voice settings. Your own transcripts with real STT errors predict better.

How many test runs do I need to compare two LLMs?

To detect a 5-point difference in task success, such as 90% versus 95%, at 80% power, plan for about 435 runs per model. Run each scenario 4 to 8 times so you can report pass^k. Smaller differences need far more runs, so focus on the metrics that actually differ.

The bottom line

The best LLM for a voice agent in October 2026 is usually a small, fast model with reasoning off, a stable cached prompt, and a stronger model routed in only for the turns that need it. Measure time to first audio, multi-turn tool accuracy and pass^k on your own calls before you commit, and measure again every time a vendor ships a new snapshot.

Related Articles