Evalgent
Back to Blog
Voice AI Evaluation

Building a Voice AI RFP: The Criteria Checklist

Deepesh Jayal
13 min read
Building a Voice AI RFP: The Criteria Checklist

The RFP is the one moment in a vendor relationship when you hold the leverage. Vendors want the deal, so they will answer what you ask — which means a weak RFP full of open-ended questions gets you a stack of marketing prose, while a strong one full of specific, evidence-demanding requirements gets you something you can actually compare. Most voice AI RFPs are weak. They ask "how accurate is your agent?" and receive "highly accurate." This checklist is how to write the version that forces real answers.

We will cover what a good RFP has to do, the criteria sections it must contain, the specific questions that expose vendors who can't back their claims, and how to structure it so responses compare. It pairs with the vendor scorecard and the bake-off playbook that come after it.

What a voice AI RFP has to do

A request for proposal has two jobs beyond gathering information. First, make responses comparable: if every vendor answers the same specific questions in the same structure, you can line the answers up side by side instead of comparing three differently shaped sales decks. Second, force evidence over claims. A vendor's self-reported "97% accuracy" means nothing without a disclosed method and your own data behind it, the problem our vendor metrics piece explains. The RFP is where you require that method up front, so a vendor either provides it or reveals that it can't.

A good RFP also protects the later stages. Every requirement you make binding here becomes a commitment you can hold the vendor to in the contract and verify in the bake-off. A vague RFP gives you nothing to enforce.

The voice AI RFP criteria checklist

Structure the RFP around these sections, and require specifics in each.

Accuracy and understanding. Ask for accuracy measured on representative calls like yours, with the exact definition and method disclosed — not a headline number. Require transcription accuracy weighted for critical entities, and intent accuracy, separately.

Latency. Require time to first audio and turn latency at p90 and p95, not averages, and under expected concurrency. Reference a real bar like the ITU-T G.114 150ms comfort limit so the answer is anchored.

Task success and escalation. Ask how the vendor measures whether calls are actually resolved, and how the agent escalates to a human, with the accuracy of that escalation.

Safety, security, and compliance. Require security attestations such as ISO/IEC 27001 or equivalent, data-handling and PII practices, and alignment with a framework like the NIST AI Risk Management Framework. Ask specifically about prompt-injection resistance and data retention.

Scalability and reliability. Require an uptime SLA, tested peak concurrency, and behavior under load, with penalties for missing the SLA.

Pricing. Require a full cost breakdown — per-minute and, where offered, per-resolution — including every pass-through fee, so the true cost is visible rather than a headline rate.

Support, roadmap, and references. Ask for support tiers, response times, the model-update roadmap, and reference customers in your industry and at your scale.

Evaluation and change management. This is the section most RFPs omit and the one that matters most: ask how the vendor tests its own agent, and exactly what happens — and how you are notified and re-validated — when it ships a model update.

The questions that expose weak vendors

A few requirements separate vendors who can back their claims from those who can't. Ask each vendor to state accuracy measured on data like yours and to disclose the method — a vendor that only has clean-audio marketing numbers will struggle. Ask what happens to your agent when they update a model, and how you are re-validated; a vague answer signals you will absorb silent regressions. Ask them to describe their own evaluation methodology in detail; if they can't, they aren't measuring rigorously. And ask them to commit to the SLA with penalties; willingness to be held accountable is itself a signal. The goal is not to trap vendors but to reward the ones operating with the rigor you need.

How to structure the RFP for comparable responses

The structure is what turns responses into a comparison.

1. Lead with mandatory requirements — List the non-negotiables (security posture, compliance, SLA floor) as pass/fail gates, so non-qualifying vendors are filtered before scoring.

2. Ask specific, uniform questions — Use the same precise questions for every vendor, so answers line up rather than sprawling in different shapes.

3. Demand evidence, not adjectives — Require measured numbers, methods, and attestations, and explicitly reject unsupported claims.

4. Weight the criteria in advance — Publish or privately fix how each section is weighted for your use case, so scoring is consistent.

5. Require your-data or bake-off validation — State that headline claims will be verified on your own calls before any award, which deters inflated answers.

6. Score responses uniformly — Evaluate every RFP against the same rubric, ideally with a neutral reviewer, to keep the comparison honest.

Weak RFP vs strong RFP

The difference is what you get back.

AspectWeak RFPStrong RFP
QuestionsOpen-endedSpecific and uniform
AnswersMarketing proseMeasured evidence
Accuracy ask"How accurate?"Measured on data like ours, method disclosed
Latency ask"Is it fast?"p90/p95 under load
ComparableNoYes, side by side
Enforceable laterNoBinding commitments

Common RFP mistakes

The errors are familiar. Asking open-ended questions that invite marketing instead of evidence. Omitting the model-update and evaluation section, so you never learn how the vendor keeps quality up. Accepting self-reported metrics without requiring the method or your-data validation. Skipping mandatory security and compliance gates, so unqualified vendors reach scoring. Leaving pricing vague, so pass-through fees surface after signing. And scoring each response differently because the questions weren't uniform. Each one lets the vendor who markets best, rather than performs best, win the paperwork — the same failure the whole independent evaluation discipline exists to prevent.

From RFP to contract

The RFP is not just a filter; it is the first draft of your contract. Every requirement you make binding — an accuracy floor, a latency ceiling, an uptime SLA, a model-update notification obligation — should carry through into the statement of work and the agreement, so what the vendor promised on paper is what you can enforce later.

This is why specificity pays off twice: it makes responses comparable now, and it gives you enforceable commitments after signing. Treat a strong RFP answer as a term, not a talking point. When the vendor's own words about accuracy, latency, and updates are written into the contract, a later shortfall becomes a breach you can act on rather than a disappointment you absorb.

Using an RFP with Evalgent

Evalgent turns RFP claims into verified facts. When vendors respond with accuracy, latency, and task-success numbers, Evalgent measures the same metrics on your own calls, so you can check every claim against reality before an award. Scenarios reproduce your real traffic and run identically against each responding vendor, Profiles vary caller conditions so no vendor passes on clean audio alone, and Metrics score each RFP criterion — accuracy, latency percentiles, escalation, safety — on your data with one fixed definition. Evaluations run at concurrency to test the reliability the SLA promises, and Reviews let your team replay any call behind a number. The RFP sets the requirements; Evalgent confirms which vendor actually meets them.

The result is an award you can defend: requirements stated up front, and responses verified on your own calls rather than taken on the vendor's word. To validate the vendors responding to your RFP on your own data, book a demo.

The bottom line

A voice AI RFP works when it forces evidence and makes responses comparable: accuracy measured on data like yours, latency at the tail under load, security attestations, SLAs, transparent pricing, and each vendor's evaluation and change-management method. Ask open questions and you get marketing; ask specific ones and you get something you can score.

Lead with mandatory gates, require methods and measured numbers, and state that claims will be verified on your own calls before any award. The RFP is your leverage moment — a strong one makes the best vendor, not the best marketer, win.

Frequently asked questions

What should a voice AI RFP include?

A voice AI RFP should include sections on accuracy and understanding, latency, task success and escalation, safety and security, scalability and reliability, pricing, support and references, and — critically — evaluation and change management. Each section should demand specific evidence rather than open-ended claims: measured numbers, disclosed methods, attestations, and SLAs. Structure the questions uniformly so responses line up side by side and can be scored consistently.

How do you write an RFP for a voice agent vendor?

Lead with mandatory pass/fail gates like security and compliance, then ask the same specific questions of every vendor, demanding measured evidence instead of adjectives. Weight the criteria for your use case in advance, and state that headline claims will be verified on your own calls before any award. Score every response against the same rubric, ideally with a neutral reviewer, to keep the comparison fair.

What questions should you ask voice AI vendors in an RFP?

Ask for accuracy measured on data like yours with the method disclosed, latency at p90 and p95 under load, how task success is measured, and escalation accuracy. Ask for security attestations and data-handling practices, an uptime SLA with penalties, a full pricing breakdown including pass-through fees, and — most revealing — how they test their agent and what happens to your deployment when they update a model.

Why do voice AI RFP responses all sound the same?

Because the questions were open-ended, inviting each vendor to answer with its most flattering marketing. "How accurate is your agent?" gets "very accurate" from everyone. Specific, evidence-demanding questions break the sameness: asking for accuracy measured on data like yours, with the method disclosed, produces answers that differ meaningfully — and reveals which vendors can actually back their claims.

What's the difference between an RFP and a bake-off?

An RFP is the document that gathers structured, comparable vendor responses and states your requirements. A bake-off is the trial that verifies those responses by running the vendors against your calls. The RFP filters and shortlists on paper; the bake-off tests the finalists in practice. Use the RFP to require evidence and the bake-off to confirm it before you commit.

How do you make RFP responses comparable across vendors?

Ask every vendor the same specific questions in the same structure, require measured evidence rather than prose, and lead with mandatory gates that filter non-qualifying vendors first. Weight and score each response against a single rubric, ideally with a neutral reviewer. Uniform questions and uniform scoring are what let you line responses up side by side instead of comparing differently shaped sales decks.

Should a voice AI RFP require security and compliance proof?

Yes, as mandatory gates. Require security attestations such as ISO/IEC 27001 or equivalent, clear data-handling and PII practices, and alignment with a recognized risk framework, and treat missing them as disqualifying rather than a scoring deduction. For regulated industries this is non-negotiable, and building it into the RFP as a gate keeps unqualified vendors from consuming evaluation effort later.

How do you verify RFP claims before signing?

Don't take numbers on the vendor's word — verify them on your own calls. Run the responding vendors through the same representative scenarios, measure the metrics they claimed with one fixed definition, and compare the results to their RFP answers. State in the RFP that claims will be validated this way, which both deters inflated responses and gives you evidence to support the award.

Related Articles