Test your voice agent
A Procurement Lead's Guide to Choosing a Voice Vendor

# A procurement lead's guide to choosing a voice agent vendor
Quick answer
A procurement lead choosing a voice agent vendor writes a clear RFP, scores every bid on the same test cases, and prices the total cost of ownership, not just the per-minute rate. Then they lock protections into the contract: SLAs, exit rights, price caps, and independent test evidence as an award condition.
A voice agent contract is not a software license. It is an operational commitment that puts an AI on the phone with your customers, in your name, at scale. When it works, it deflects volume and cuts cost per call. When it fails, it fails one caller at a time, and the exposure lands on your desk during the renewal.
Most of the buying advice online is written for engineers or product owners. This guide is written for the procurement lead. Your job is not to admire the technology. Your job is to source it fairly, price it honestly, and paper it so the company is protected when the demo magic wears off. That means a defensible request for proposal, a scoring method that compares like with like, and contract terms that survive contact with production.
This post covers writing the RFP, scoring bids apples-to-apples, the total cost of ownership beyond the per-minute rate, the contract terms that matter, third-party risk, and how to require independent test evidence as an award condition. The technical depth belongs to your CTO and engineering leads. The structure of the deal belongs to you.
Why voice agent selection lands on procurement
Voice agents get bought under pressure. A business unit runs a slick pilot, falls for a demo, and arrives at your door with a preferred vendor already chosen. The technical champion is enthusiastic. The timeline is short. The pressure is to paper the deal and move on.
That is exactly the moment a procurement lead earns their keep. A demo is a curated success under ideal conditions. It tells you almost nothing about how the agent handles the callers you actually get: the accents, the interruptions, the angry ones, the edge cases. Sole-sourcing off a demo leaves you with no comparable evidence and no leverage in the negotiation.
The discipline you already apply to any strategic supplier applies here. Public buyers formalize it in the Federal Acquisition Regulation at acquisition.gov, and the same principles scale down to any enterprise: competition, documented criteria, and an award you can justify. A voice agent is a new category of purchase, but it is not a new kind of decision. Treat it like the high-risk, high-visibility supplier relationship it is.
For the wider evaluation picture that your technical peers care about, our pillar on how to evaluate voice agent vendors lays out the scorecard. This guide sits alongside role-specific views for the CTO, the VP of engineering, the CX leader, the compliance officer, and the contact center manager.
Writing an RFP that produces comparable bids
The point of an RFP is not paperwork. It is comparability. If every vendor answers a different question, you cannot score them side by side, and the loudest sales team wins by default.
Write the RFP so that answers snap into a grid. Ask closed questions where you can. Require structured responses. Forbid marketing prose in the technical sections. A good RFP forces vendors onto your terms rather than theirs.
Cover these areas at minimum:
- Functional scope. The use cases, call types, languages, and integrations you need. Be specific about volume and peak concurrency.
- Performance evidence. How the vendor measures accuracy, task success, containment, and latency, and whether they will submit to independent testing on your scenarios.
- Pricing model. The full rate card, not a headline per-minute number. Ask for every line item, minimum commitments, and overage rates.
- Security and compliance. Certifications, data handling, subprocessors, and where recordings and transcripts live.
- Service commitments. Uptime, support tiers, incident response, and change-management practices.
- Contract posture. Willingness to accept your paper, term length, and exit terms.
Do not let vendors grade their own homework. Any metric a vendor self-reports was measured on data they chose, under conditions they controlled. Require that performance claims be verifiable on your test cases. The difference between checking claims yourself and taking them on faith is the difference explained in our guide on testing vs evaluation for voice agents.
Scoring bids apples-to-apples
A weighted scorecard is the backbone of a defensible award. Set the criteria and the weights before any bid arrives. If you invent the weights after you have seen the responses, you are rationalizing a favorite, not evaluating a field.
Weight the criteria by what matters to your business. Accuracy and task success usually carry the most weight, followed by cost, security, and service commitments. Write the weights down and freeze them. Circulate them to the evaluation committee so no one can move the goalposts later.
The hardest part is the performance section, because vendor-reported numbers do not compare. One vendor's "95% accuracy" and another's "97%" were measured on different audio, different tasks, and different definitions. The only fair number is one you produce yourself, on the same test cases, scored the same way for every vendor. Our guide to compare voice agents on the same test cases walks through the mechanics.
This is where independence matters most. When a neutral party runs your scenarios across every bid and scores them identically, the performance column of your scorecard becomes evidence rather than assertion. That is the premise of independent voice AI evaluation, and it is the single strongest input a procurement lead can put in front of an award committee.
Total cost of ownership beyond the per-minute rate
The per-minute rate is the number vendors want you to anchor on. It is rarely the number that matters. Total cost of ownership is what you actually pay to run the agent over the life of the contract, and it hides in the line items the headline rate leaves out.
Build a TCO model before you negotiate. Include:
- Usage charges. Per-minute or per-session rates, minimums, and overage tiers. Model your real volume, including peaks.
- Integration and build cost. Connectors to your telephony, CRM, and knowledge base. One-time and ongoing.
- Human handoff cost. Every call the agent fails to contain still costs a live agent. A cheaper agent that escalates more can cost more overall. Our guide to containment vs deflection explains why that distinction moves the number.
- Testing and monitoring. The cost of proving the agent works before launch and watching it after.
- Switching cost. What it takes to leave, which is the hidden tax of vendor lock-in.
Then compare pricing models on their merits. A per-minute deal rewards the vendor for long calls. An outcome-based model can align the vendor with your goals but needs a trusted measure of the outcome. Our breakdown of outcome-based voice agent pricing covers the trade-offs. The right model depends on your volume profile and your appetite for shared risk.
Whatever the model, price the whole relationship. The cheapest quote and the lowest total cost are often different vendors.
Contract terms that protect you
The contract is where procurement leaves its mark. A voice agent relationship usually sits under a master service agreement with order forms beneath it. Get the following into the paper, not the sales email.
Service level agreements. Define uptime, latency, and support response as numbers with credits attached. A promise without a remedy is decoration.
Exit and portability. Spell out how you leave: notice period, data export in a usable format, transition assistance, and deletion. The time to negotiate an exit is before you sign, when you still have leverage.
Price protection. Cap annual increases. Lock overage rates for the term. Protect yourself against a renewal that resets the economics you modeled.
Model-update notice. Voice agents change under you. The vendor can swap the underlying model or tune the prompts, and behavior shifts overnight. Require advance notice of material changes and the right to re-test before they reach production. Without this clause, your evaluation has a shelf life of one release.
Performance as a condition, not a hope. Tie acceptance and, where possible, renewal to measured performance on your test cases. Make the evidence a contractual deliverable.
Liability, indemnity, and compliance. Standard for any strategic supplier, and non-negotiable when an AI is speaking to customers in regulated contexts.
Managing third-party and vendor risk
A voice agent vendor is a third party handling your customers and, often, their personal data. That makes it a live entry in your vendor-risk program, not a one-time procurement event.
Do the diligence you would for any critical supplier. Confirm security posture with real evidence, not a logo on a slide. A SOC 2 report tells you an independent auditor examined the vendor's controls; ask for the report, not just the claim. Map the subprocessors and where data lives. Get a data processing agreement that matches your obligations.
Frame the AI-specific risks against a recognized standard. The NIST AI Risk Management Framework gives you a shared vocabulary for the failure modes that matter: an agent that gives wrong answers, leaks information, or acts outside policy. Requiring vendors to speak to those risks separates the mature suppliers from the ones still selling a demo.
Concentration risk deserves a mention. Betting everything on one vendor raises the cost of switching and the pain of an outage. Many enterprises deliberately keep options open. The reasoning is laid out in our pillar on voice agent evaluation, and it is a legitimate procurement position, not a technical preference.
How to run a procurement evaluation for a voice agent vendor
Follow a repeatable process so the award holds up when someone asks why you chose this vendor. Each step produces a record.
1. Define requirements and weights. Gather functional, performance, security, and commercial needs from the business and technical stakeholders. Turn them into a weighted scorecard, and freeze the weights before bids arrive.
2. Issue the RFP. Send a structured RFP that forces comparable answers. Set a clear deadline and a single channel for questions so every vendor sees the same information.
3. Shortlist on paper. Score the written responses against your frozen criteria. Cut vendors who cannot meet a hard requirement before you spend time testing.
4. Build one shared test set. Assemble real scenarios from your own call traffic, including the edge cases that break agents. This same set goes to every shortlisted vendor.
5. Run independent testing. Have a neutral party run the shared test set across every vendor and score the results identically. This fills the performance column with evidence, not vendor claims.
6. Model total cost of ownership. Price each vendor over the full term, including handoff, integration, and switching costs, not just the per-minute rate.
7. Negotiate the contract. Secure SLAs, exit rights, price protection, model-update notice, and performance as an award condition. Negotiate exit before you sign.
8. Award and document. Combine the scores, TCO, and risk findings into a written rationale. Keep the criteria, test results, and reasoning so the decision is auditable.
Procurement criteria, contract terms, and red flags
Use this table as a working checklist. The left column is what you evaluate. The middle is what to require in the contract. The right is the signal that should slow the deal down.
| Procurement criterion | What to require in the contract | Red flag |
|---|---|---|
| Performance | Acceptance and renewal tied to measured results on your test cases | Only self-reported metrics; refuses independent testing |
| Pricing | Full rate card, capped increases, locked overage rates | Headline per-minute rate with vague "custom" line items |
| Total cost of ownership | Modeled over the full term, including handoff and switching | Quote covers usage only; switching cost undiscussed |
| Service levels | Uptime, latency, and support as numbers with credits | Promises without remedies or measurable thresholds |
| Exit and portability | Data export, transition help, deletion, defined notice | No exit terms; proprietary formats that trap your data |
| Model changes | Advance notice and a right to re-test before production | Vendor can change the model silently at any time |
| Security and privacy | Current SOC 2, data processing agreement, subprocessor list | Certification claimed but report never produced |
| Vendor risk | Fits your third-party risk program and NIST-aligned review | Single point of failure with no continuity plan |
Where independent evaluation fits the procurement process
The recurring theme above is evidence. Every strong procurement position depends on a performance number you can trust, and vendor-reported numbers do not qualify.
This is the role Evalgent is built for. Evalgent is an independent, third-party evaluation platform for AI voice agents. It runs your test cases across every vendor, scores the results with no stake in which vendor wins, and hands you the comparable evidence your scorecard needs. As a procurement lead, you can make that independent evaluation a formal award condition and a contract gate: no vendor advances, and no renewal clears, without passing the same test on the same terms.
That turns "the demo was impressive" into a documented, defensible decision. It also gives you leverage in the negotiation, because a vendor confident in their product will accept independent measurement, and a vendor that refuses has told you something. For the ongoing version of this, see our guide to the third-party voice agent audit, which extends the same evidence discipline past the award and into the life of the contract.
Frequently asked questions
How does a procurement lead choose a voice agent vendor?
Start with a weighted scorecard and a structured RFP so bids compare directly. Shortlist on paper, then run one shared test set across every vendor through an independent evaluator. Model total cost of ownership over the full term. Award on the combined evidence, and lock protections into the contract before signing.
What should go into a voice agent vendor RFP?
Cover functional scope, performance evidence, the full pricing model, security and compliance, service commitments, and contract posture. Ask closed, structured questions so answers snap into a scoring grid. Require that performance claims be verifiable on your own test cases rather than accepted from vendor marketing materials.
How do you compare voice agent vendor bids fairly?
Set weighted criteria before bids arrive and freeze them. Score written responses against those criteria. For performance, ignore self-reported numbers and instead run one shared test set across every vendor, scored identically by a neutral party. That produces apples-to-apples evidence instead of comparing marketing claims measured under different conditions.
What is the total cost of ownership of a voice agent?
It is everything you pay over the contract, not the per-minute rate alone. Include usage charges, integration and build cost, the cost of calls escalated to humans, testing and monitoring, and the switching cost of leaving. The cheapest quote and the lowest total cost of ownership are often different vendors.
What contract terms should procurement require from a voice agent vendor?
Require service level agreements with credits, exit and data-portability rights, price protection with capped increases, and advance notice before model changes with a right to re-test. Tie acceptance and renewal to measured performance on your test cases. Negotiate exit terms before signing, while you still hold leverage.
Why require independent test evidence as an award condition?
Vendor-reported metrics are measured on data the vendor chose under conditions they controlled, so they do not compare. Making independent evaluation an award condition puts comparable, neutral evidence in your scorecard. It also signals confidence: a strong vendor accepts measurement, and one that refuses has revealed a concern worth investigating.
How do you manage third-party risk for a voice agent vendor?
Treat the vendor as a critical supplier in your vendor-risk program. Confirm security with a real SOC 2 report, a data processing agreement, and a subprocessor list. Frame AI-specific risks against the NIST AI Risk Management Framework, and consider concentration risk before betting entirely on a single provider.
What are red flags in a voice agent vendor contract?
Watch for vendors who only offer self-reported metrics or refuse independent testing, vague "custom" pricing with no rate card, no exit or data-export terms, proprietary formats that trap your data, promises without measurable remedies, and the right to change the underlying model silently. Each shifts risk onto you after signing.
The bottom line
A defensible voice agent award rests on comparable evidence, a full total-cost model, and a contract that protects you after the demo ends. Make independent evaluation an award condition, and the decision holds up long after the sales cycle closes.
Ready to make independent evidence your contract gate? Book a demo to see how Evalgent scores every vendor on the same test cases, so your next award is one you can defend.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more