Evalgent
Back to Blog
Voice AI Evaluation

Voice Agent SLA: What to Require from a Vendor

Deepesh Jayal
12 min read
Voice Agent SLA: What to Require from a Vendor

# Voice Agent SLA: What to Require from a Vendor

> Quick answer: Voice agent SLA requirements are the guarantees a vendor commits to in writing. They cover uptime, P95 and P99 latency ceilings, containment floors, escalation, model-update notice, incident response, and service credits. Each clause needs a defined measurement method and independent verification to hold value.

A service-level agreement is only as strong as its numbers. Most voice agent contracts promise "high availability" and "low latency." Those words do not survive a bad week. When calls drop or the agent stalls, you need a clause you can point to.

This guide lists the clauses that belong in a voice AI SLA. It explains the target for each one. It also shows how each metric is measured. And it shows how an independent party verifies it. The goal is a contract you can enforce, not a brochure.

What a voice agent SLA actually is

> Voice agent SLA: a written contract that defines the service levels a voice AI vendor guarantees. It pairs each promise with a target, a measurement window, and a remedy when the target is missed.

An SLA is different from an SLO. A service-level objective is an internal goal. An SLA is a customer-facing promise with consequences. Frameworks like ITIL and ISO/IEC 20000 treat both as core to service management.

Voice agents raise the stakes. A web API that slows down annoys users. A voice agent that stalls mid-sentence loses the call. The caller hears silence and hangs up. So a voice SLA must cover more than server uptime. It must cover the live conversation itself.

Buyers often accept the vendor's default template. That is a mistake. The default protects the vendor. Your job is to add the clauses below. Then set targets you can defend. For a broader view, see our guide on how to evaluate voice agent vendors.

The clauses that belong in a voice AI SLA

Each clause below solves a specific failure mode. Skip one, and that failure has no remedy.

Uptime and availability

Uptime is the share of time the service is reachable. It is usually stated as a percentage, such as 99.9%. That figure sounds solid. It still allows about 8.8 hours of downtime a year.

Push for 99.9% at a minimum. Ask for 99.95% on core call paths. Then read the fine print. Many SLAs exclude scheduled maintenance and third-party outages. Those carve-outs can swallow the guarantee. Define availability by successful call handling, not just a server ping.

P95 and P99 latency ceilings

Average latency hides the worst calls. A vendor can post a great average. Yet one call in twenty can feel broken. That is why you cap percentiles instead.

The P95 value is the latency) that 95% of turns beat. The P99 is the level that 99% beat. Set a P95 ceiling for typical experience. Set a P99 ceiling for the tail. A common target is P95 under 1.5 seconds for response time. Pair it with P99 under 2.5 seconds.

Be precise about what you measure. Latency and response time are not the same thing. Our guide on latency for voice agents breaks down each segment of the turn.

Accuracy and containment floors

Speed means little if the agent gets things wrong. Set an accuracy floor for intent recognition. Set a separate floor for correct actions taken.

Containment is the share of calls the agent resolves without a human. Set a containment floor tied to real resolution. Do not tie it to simple call deflection. The two are different. Vendors sometimes blur them. Our guide on containment versus deflection explains the trap.

Escalation guarantees

Some calls should reach a human. A refund dispute or a safety issue cannot wait. Your SLA needs an escalation clause with a time target.

Require the agent to hand off within a set number of seconds once a trigger fires. Require context to travel with the call. The human should not start from zero. See our guide on escalation for voice agents for trigger design.

Model-update notice

Vendors update models often. A silent update can change agent behavior overnight. A script that passed on Monday may fail on Tuesday.

Require written notice before any model or prompt change on your account. Ask for a rollback path. Ask for a test window so you can re-run your own checks first. This clause protects the work you did during evaluation.

Incident response and communication

Outages happen. What matters is the response. The NIST incident handling guide offers a solid model for structure and timing.

Require a response-time commitment by severity. A total outage needs faster acknowledgment than a minor bug. Require a status page and a named contact. Require a post-incident report within a set number of days.

Remedies and service credits

A promise with no penalty is a suggestion. Service credits give the SLA teeth. They refund a share of fees when the vendor misses a target.

Push for credits that scale with severity. A brief blip and a full-day outage should not cost the same. Confirm how credits are claimed. Some contracts require you to file within a short window or forfeit the credit.

How each clause is measured and verified

The table below pairs each clause with a realistic target. It also shows the way to confirm it. Treat the targets as starting points. Then adjust for your call volume and risk. This one table is the core of an enforceable voice agent SLA.

SLA clauseTarget to requireHow to measure and verify
Uptime and availability99.9%+ on core call pathsIndependent synthetic calls plus vendor logs, reconciled monthly
P95 latencyUnder 1.5s response timeTimestamped turn logs, percentile computed by a third party
P99 latencyUnder 2.5s response timeTail sampled from full call set, not a curated subset
Intent accuracy90%+ on your top intentsBlind grading of transcripts against a labeled test set
ContainmentFloor tied to real resolutionOutcome audit that separates resolution from deflection
Escalation timeHandoff within 10s of triggerTrigger-to-human timestamps on sampled escalation calls
Model-update noticeWritten notice before any changeChange log with dates, checked against observed behavior
Incident responseAcknowledge by severity tierTimestamped ticket history versus the committed window
Service creditsScaled to severity and durationIndependent breach count matched to the credit schedule

The right column matters most. A number you cannot check is not a guarantee. It is a hope. This is where an independent party earns its place. We return to that point below.

How to negotiate and instrument a voice agent SLA

Follow these steps to turn a vendor template into a contract you can enforce. Do the measurement work before you sign, not after.

1. List your failure modes first. Write down what a bad call looks like for your business. Map each failure to a clause above. This gives you a checklist for the vendor's draft.

2. Set targets from your own data. Use your call volume and your intents. Do not accept generic numbers. Benchmark the agent on your test cases first, as covered in our voice agent evaluation guide.

3. Define every measurement window. State whether uptime is monthly or quarterly. State the sample size for latency percentiles. Vague windows favor the vendor.

4. Separate resolution from deflection. Write the containment definition into the contract. Tie credits to real outcomes. Do not tie them to calls that simply ended without a human.

5. Add the model-update clause. Require notice and a test window. Protect your evaluation work from silent changes.

6. Name the verifier. Decide who confirms each metric. Reserve the right to an independent voice AI evaluation rather than trusting vendor dashboards alone.

7. Instrument before launch. Set up synthetic calls and transcript grading in advance. You want a baseline on day one, not a scramble after the first outage.

8. Schedule the reconciliation. Agree on a monthly review of vendor logs against independent measurements. Put the meeting in the contract.

Why independent verification changes the contract

A vendor grading its own SLA has a conflict. The party that owes credits also counts the breaches. Even honest vendors pick favorable windows and curated samples. That is human nature, not fraud.

Independent verification removes the conflict. A third party runs its own synthetic calls. It computes percentiles from the full call set. It grades transcripts blind against a labeled test set. Then it reconciles those numbers with the vendor's logs.

Evalgent works as that neutral party. It measures uptime, latency percentiles, containment, and escalation on your own scenarios. It reports what actually happened, not what a dashboard claims. When a breach occurs, you have evidence, not a debate.

This is the difference between an SLA that looks good and one that pays out. For the wider case, see our guide on the third-party voice agent audit. Verification is what makes every target above real.

Common mistakes buyers make

Buyers accept averages instead of percentiles. An average hides the tail. The tail is where callers give up. Always cap P95 and P99.

Buyers skip the containment definition. They celebrate a high number that only counts ended calls. Write resolution into the contract.

Buyers forget the model-update clause. Then behavior shifts, and no one knows why. A notice requirement prevents the surprise.

Buyers trust the vendor's own reporting. That reporting is not neutral. Independent checks keep everyone honest.

Frequently asked questions

What should a voice agent SLA include?

A voice agent SLA should include uptime, P95 and P99 latency ceilings, accuracy and containment floors, escalation time targets, model-update notice, incident response commitments, and service credits. Each clause needs a defined measurement window and a named party to verify it. Without measurement rules, the promises cannot be enforced.

What uptime should a voice agent SLA guarantee?

A voice agent SLA should guarantee at least 99.9% uptime, and 99.95% on core call paths. Define availability by successful call handling, not just a server ping. Check the exclusions carefully. Scheduled maintenance and third-party carve-outs can quietly erase the guarantee you thought you had.

How do you measure voice agent latency in an SLA?

Measure voice agent latency with timestamped turn logs from the full call set. Compute percentiles rather than averages. Cap P95 for typical experience and P99 for the tail. A third party should compute these values from unfiltered data. A curated sample can make a slow agent look fast.

What is a good containment rate for a voice agent?

A good containment rate depends on the use case, so require a floor tied to real resolution. Simple tasks support higher containment than complex disputes. Make sure the number counts calls the agent actually resolved. Deflection alone, where the call simply ends, is not the same as containment.

How are voice agent SLA service credits calculated?

Service credits are usually a percentage of fees refunded when the vendor misses a target. Push for credits that scale with severity and duration. A brief blip should cost less than a full outage. Confirm the claim process too, since some contracts require you to file within a short window.

Can you verify a voice agent SLA independently?

Yes, a voice agent SLA can be verified independently. A neutral party runs its own synthetic calls, computes latency percentiles from full data, and grades transcripts blind. It then reconciles those results with vendor logs. Independent verification removes the conflict of a vendor grading its own performance and gives you real evidence.

What is the difference between an SLA and an SLO?

An SLA is a customer-facing contract with remedies when targets are missed. An SLO is an internal objective a team sets for itself. The SLA carries consequences like service credits. The SLO guides engineering work but does not obligate the vendor to you. Buyers should negotiate the SLA, not the SLO.

Who measures voice agent SLA compliance?

Compliance is often measured by the vendor, which creates a conflict of interest. The party that owes credits should not be the only one counting breaches. An independent evaluator solves this. It measures uptime, latency, and containment on your scenarios. Then it compares its findings against vendor logs during a scheduled monthly review.

The bottom line

A voice agent SLA is only as strong as its measurement rules and its remedies. Independent verification is what turns each target into a guarantee you can enforce.

Set your targets from your own data, define every window, and name a neutral party to check the numbers. Book a demo to see how Evalgent verifies voice agent SLA metrics on your own scenarios.

Related Articles