Test your voice agent
Hidden Costs in Voice Agent Pricing You Won't See on the Rate Card

Every voice agent vendor leads with a headline number. A few cents per minute, maybe less at volume. It looks cheap, it fits neatly in a spreadsheet, and it makes the buying decision feel easy. Then the first real invoice arrives, and the finance team asks why the bill is three or four times the model they built. The gap is not fraud. It is structure. The advertised rate covers a narrow slice of what actually happens on a call, and the rest arrives as separate line items — or worse, as costs that never appear on any invoice at all.
This post maps the gap between the advertised rate and the true all-in cost. It is not a general per-minute explainer, and it is not about cost per resolution as a metric. It is about the surprises — the fees, the engineering, and the quality failures that quietly inflate the real bill.
Why the advertised rate is only a fraction
A per-minute rate is a unit cost, and unit costs are seductive because they scale linearly in your head. Ten thousand minutes at five cents feels like five hundred dollars. Clean, predictable, done.
The trouble is that a voice agent is not one product. It is a stack of metered services stitched together, plus the human and engineering effort to keep it running. The advertised rate usually reflects the vendor's orchestration layer and little else. Everything underneath it — and everything around it — is priced separately or absorbed by your own team. Proper total cost of ownership accounting counts all of it, not just the sticker.
The result is a familiar pattern. The rate you compare across vendors is the one number that varies least. The costs that actually determine your bill are the ones nobody puts on the slide.
The pass-through component fees
Under the hood, a voice agent runs a pipeline. Speech-to-text turns the caller's audio into words. A language model decides what to say. Text-to-speech turns the reply back into audio. Telephony carries the call. Each of these is a metered service, and each is often billed separately from the orchestration rate.
Here is where buyers get surprised. Some vendors quote a rate that bundles these components; others quote a thin orchestration fee and pass the component costs straight through to you. Two vendors can advertise nearly identical headline rates while one bill lands at double the other, purely because of what the number includes.
Watch these four in particular:
- Speech-to-text. Priced per minute of audio processed, sometimes with premiums for higher accuracy models or specialized vocabularies.
- The language model. Usually priced per token. A chatty agent, a long system prompt, or heavy retrieval context all raise token consumption per call — invisibly.
- Text-to-speech. Priced per character or per minute of generated audio. Premium, more natural voices cost meaningfully more than basic ones.
- Telephony. Per-minute carrier charges, plus per-number rental, and sometimes separate inbound and outbound rates.
None of these is hidden in a sinister sense. They are simply below the fold. The advertised rate answers "what does the orchestration cost," when the question you need answered is "what does a completed call cost."
Integration and engineering cost
The rate card assumes the agent already works with your systems. It never does on day one.
A voice agent that cannot look up an order, check an appointment, or authenticate a caller is a very expensive answering machine. Making it useful means integration work: connecting your CRM, your scheduling system, your knowledge base, your telephony routing. That work is engineering time, and engineering time is money the vendor's rate never mentions.
This cost is a classic transaction cost — the effort of making the deal actually function, over and above the price of the thing itself. It shows up as weeks of developer effort, ongoing maintenance when your backend APIs change, and the internal coordination to test and deploy safely. For a serious deployment it can dwarf the first year of usage fees. Our guide on why voice agents fail in production covers how thin integration quietly sinks otherwise promising pilots.
Build this into your model as a real line, not a footnote. A rate that looks great per minute can still lose to a slightly pricier vendor whose integrations are turnkey.
The cost of failed and unresolved calls
Every call you pay for is not a call that worked. This is the line item buyers understand least, and it is often the largest.
Consider what happens when an agent takes a call, spends ninety seconds fumbling, and fails to resolve the issue. You paid for those ninety seconds. The caller's problem still exists. So they call back — and you pay again. Or they escalate to a human, and you pay for the agent minutes and the human minutes. A single unresolved issue can generate two, three, or four billed interactions before it closes.
This is why the per-minute rate is a misleading unit. The unit that matters is a resolved contact, which is exactly the argument in our cost-per-resolution breakdown. If half your calls fail, your true cost per resolution is double the arithmetic — before you count the repeat calls the failures spawned.
Failed calls also carry a downstream cost that never touches the invoice. A caller who could not get help is a caller more likely to churn. That lost revenue is real, and it belongs in any honest cost accounting of the deployment.
Escalation and human-handoff cost
Handoff to a human is the safety net, and it is also a cost multiplier. When the agent gives up, the call routes to a person — and now you are paying for both.
The economics only work if the agent contains a high share of calls on its own. A deployment that escalates most calls has essentially added an expensive front-end to your existing contact center rather than replacing any of it. Worse, poorly designed handoffs frustrate callers, who must repeat themselves to the human, extending the human's handle time and cost.
There is a subtle trap here too. Vendors sometimes report high containment while quietly counting deflection — calls the agent ended without truly resolving — as contained. That inflates the apparent savings. The distinction matters enough that we wrote a full explainer on containment versus deflection. A call the agent hung up on is not a call it handled.
Overage and volume-tier surprises
Pricing tiers are where the model you built and the bill you receive diverge most sharply.
Volume discounts sound like pure upside. In practice they come with commitments, minimums, and overage rates that reset the math. A few patterns to watch:
- Minimum commitments. You agree to a monthly floor. If your volume dips — a slow season, a product change — you pay for minutes you never used.
- Overage penalties. Exceed the tier and the marginal rate can jump well above your blended average, so a good month costs disproportionately more.
- Feature gating. The advertised rate may exclude analytics, recordings, or premium voices that live in a higher tier you end up needing.
- Annual true-ups. Some contracts reconcile usage annually, producing a surprise charge long after the budgeting window closed.
Tiered pricing rewards predictable, steady volume. Real contact volume is spiky. Model your best and worst months, not just the average, or the overage line will find you.
The hidden cost of low quality
The most expensive line item is the one with no line at all: poor quality.
A voice agent that mishears names, loses context mid-conversation, or answers confidently but wrongly does not just fail the call it is on. It erodes trust in every future call. Customers who have one bad experience avoid the channel, call during peak hours, demand a human immediately, or leave for a competitor. None of that shows up on the vendor invoice, and all of it hits your P&L.
Latency is part of this. Long silences and awkward pauses make callers hang up or talk over the agent, spiking failure rates. The ITU-T G.114 recommendation on acceptable one-way latency for voice is a useful reference point for how sensitive people are to delay. An agent that feels sluggish is an agent that loses calls.
Quality is the multiplier that sits behind every other cost in this post. High quality means more resolutions, fewer repeats, fewer escalations, and fewer lost customers. Low quality inflates every one of them at once. This is why measuring quality before you sign — not after the bill arrives — is the single highest-leverage thing a buyer can do.
Advertised versus all-in: a worked comparison
The table below shows how a clean advertised rate expands once every real line item is counted. The figures are illustrative ballparks, not vendor quotes — the point is the shape of the gap, not the exact numbers.
| Cost line | On the rate card? | Illustrative impact on per-minute cost |
|---|---|---|
| Orchestration (the advertised rate) | Yes | ~$0.05 |
| Speech-to-text | Sometimes | +$0.01–$0.03 |
| Language model tokens | Rarely | +$0.02–$0.06 |
| Text-to-speech | Sometimes | +$0.01–$0.03 |
| Telephony | Rarely | +$0.01–$0.02 |
| Integration and maintenance | No | Amortized; often large upfront |
| Failed and repeat calls | No | Multiplies cost per resolution |
| Human escalation | No | Adds agent minutes at full loaded cost |
| Overage and tier penalties | Buried in contract | Spikes in peak months |
| Lost customers from poor quality | No | Off-invoice, potentially the largest |
| All-in effective rate | — | Often $0.20+ per minute |
An advertised rate near five cents can land above twenty cents all-in once the pass-through components alone are stacked — and that is before failed calls and escalations bend the cost-per-resolution number upward.
How to uncover the true all-in cost
Follow these steps to replace the advertised rate with a number you can actually budget against.
1. List every metered component. Ask the vendor, in writing, whether speech-to-text, the language model, text-to-speech, and telephony are bundled or passed through. Get per-unit rates for each.
2. Estimate consumption per call. Multiply typical call length and token usage by your real volume. Model an average call and a long, messy one — the messy ones drive the bill.
3. Quantify integration up front. Scope the engineering days to connect your systems and maintain them. Amortize that across expected call volume as a real per-minute add.
4. Measure resolution rate, not call count. Run representative calls and count how many actually close the issue. Divide total cost by resolutions, not by minutes. Our vendor evaluation guide walks through this.
5. Price the escalations. Estimate the share of calls handed to humans and add the loaded human cost for those minutes on top of the agent minutes.
6. Stress-test the tiers. Model your highest and lowest volume months against the contract's minimums and overage rates, not just the average.
7. Put a number on quality risk. Estimate repeat-call rate and likely churn from poor experiences. Even a rough figure keeps the largest hidden cost visible in the decision.
Work through all seven and the all-in cost of a voice agent stops being a mystery. The procurement checklist turns these questions into contract language before you sign.
Frequently asked questions
Why is my voice agent bill higher than the advertised per-minute rate?
Because the advertised rate usually covers only the orchestration layer. Speech-to-text, the language model, text-to-speech, and telephony are often billed separately or passed through. Add integration, escalations, and overage, and the effective rate commonly lands several times higher than the headline number you compared across vendors.
What are pass-through component fees in voice agent pricing?
Pass-through fees are the costs of the underlying services a vendor charges you directly rather than absorbing. The four main ones are speech-to-text, language model tokens, text-to-speech, and telephony. Some vendors bundle them into one rate; others quote a thin orchestration fee and forward each component cost to your invoice.
How much does integration add to voice agent costs?
Integration is engineering time to connect your CRM, scheduling, knowledge base, and telephony, plus ongoing maintenance. It never appears on a rate card. For a serious deployment it can exceed the first year of usage fees, so treat it as a real line item amortized across your expected call volume.
Do I pay for failed voice agent calls?
Yes. You are billed for the minutes of a call whether or not it resolves the issue. A failed call often triggers a repeat call or a human escalation, so a single unresolved issue can generate several billed interactions. This is why cost per resolution matters more than cost per minute.
What is the difference between containment and deflection in pricing?
Containment means the agent genuinely resolved the call without a human. Deflection means the agent ended the call without necessarily solving anything. Some vendors count both as contained, inflating apparent savings. A deflected caller who calls back or churns costs you more, not less, than an honest escalation would have.
How do overage and volume tiers create surprise costs?
Volume tiers add minimum commitments, overage penalties, and feature gating. You pay a floor even in slow months, and exceeding a tier can spike the marginal rate above your blended average. Some contracts also reconcile usage annually, producing a charge long after the budget window closed. Model peak and trough months, not just the average.
What is the hidden cost of low voice agent quality?
Low quality drives repeat calls, escalations, and customer churn — none of which appear on the vendor invoice. A caller who has a bad experience avoids the channel or leaves for a competitor. Because quality multiplies every other cost, measuring it before signing is the highest-leverage move a buyer can make.
How do I compare voice agent vendors on true cost?
Do not compare advertised rates. Compare all-in cost per resolution: metered components, integration, escalations, and quality risk included. Run the same representative calls against each vendor, measure real resolution rates, and normalize everything to a resolved contact. The cheapest headline rate frequently loses on the number that actually determines your bill.
Get the true cost with Evalgent
Evalgent is an independent testing and evaluation platform that surfaces the quality and resolution reality that ultimately determines your true all-in cost. It works through five primitives:
- Scenarios — reproducible call situations, including the messy edge cases that inflate token usage and failure rates.
- Profiles — realistic caller personas so your resolution numbers reflect real people, not scripts.
- Metrics — resolution, containment, latency, and accuracy, defined consistently across every vendor.
- Evaluations — automated scoring of whether each call actually closed the issue, so cost per resolution is real, not assumed.
- Reviews — human verification of borderline calls, so deflection never gets counted as containment.
Because Evalgent measures what really happens on a call, you can price vendors on all-in reality before you sign. Ready to see the true cost behind the rate card? book a demo.
The bottom line
The advertised per-minute rate is the least useful number in voice agent pricing. What determines your bill is everything underneath and around it.
Count the pass-through components, the integration work, the failed and escalated calls, the tier penalties, and — above all — the quality that multiplies every other cost. An advertised nickel can become a quarter all-in. Measure resolution and quality before you sign, and the surprises stop being surprises.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more