Evalgent
Back to Blog
Voice AI Evaluation

How to Evaluate a Retail Support Voice Agent Vendor

Deepesh Jayal
12 min read
How to Evaluate a Retail Support Voice Agent Vendor

# How to evaluate a retail support voice agent vendor

Quick answer

To evaluate retail support voice agent vendor options, run your own catalog, orders, and peak-season traffic through each one. Score order lookup, returns and exchanges, refunds, catalog grounding, tool calls, containment, CSAT, and cost per resolution. Decide on measured results under real load, not a scripted holiday-season demo.

Retail support is where voice agents look best in a demo and worst in production. A vendor's sales call uses a clean catalog and one polite caller. Your December looks nothing like that. Order volumes spike, callers are angry, and the same agent now answers thousands of calls at once.

This guide shows how to evaluate retail support voice agent vendor options on the traffic that actually breaks agents. It covers the six dimensions that matter, the pass bars to hold each vendor to, and a step-by-step process. Choosing an ecommerce support voice AI vendor is a recurring, high-stakes call. It builds on our voice agent vendor scorecard, applied to ecommerce support specifically.

Why retail support breaks voice agents differently

Retail is not generic support. It combines high volume, seasonal spikes, and money changing hands. Each of those raises the stakes.

Volume is the first difference. A retail line handles order lookup, returns and exchanges, refunds, and shipping and tracking questions all day. The most common request is WISMO — "where is my order." It sounds simple. It hides a chain of tool calls to your order system.

Seasonality is the second. Black Friday and Cyber Monday can multiply your call volume overnight. Concurrency climbs from hundreds of calls to tens of thousands. Latency that felt fine in October falls apart under peak load.

Money is the third. A refund is a financial action. A wrong promo code costs margin. A hallucinated price or stock count creates a promise you cannot keep. That is why grounding on your real catalog matters more here than almost anywhere else.

For the underlying metrics and target ranges, see our customer support voice agent metrics guide. This post focuses on the vendor decision.

The six evaluation dimensions that matter

A retail voice agent evaluation should score six things. Read them together, not in isolation. A vendor can look strong on catalog grounding and still fail under peak load.

The table below is your scoring frame. For each dimension, it gives what to test, the pass bar to hold, and the red flag that should end the conversation. Set your own thresholds where your risk tolerance differs.

DimensionWhat to testPass barRed flag
Order, return, and refund handlingOrder lookup, returns within and outside the window, exchanges, refund statusCorrect action on 95%+ of scripted order scenariosIssues a refund outside policy, or invents an order status
Catalog and policy groundingProduct, stock, price, and returns-window questions on your live catalogZero hallucinated stock, price, or policy across the test setStates a price, stock level, or policy not in your catalog
Tool-call correctnessCalls to order, inventory, and loyalty systems with the right arguments98%+ correct tool selection and arguments on valid inputsCalls the wrong system, or fabricates a result when a call fails
Peak-load performanceConcurrency ramps to your Black Friday peak, sustained for an hourLatency and accuracy hold within 15% of low-load numbersLatency doubles, calls drop, or accuracy falls under load
Containment and escalationAngry callers, edge cases, and requests the agent cannot solveEscalates cleanly with full context on 100% of out-of-scope callsTraps callers, or escalates with no context handed to the agent
CSAT and cost per resolutionPost-call sentiment plus fully loaded cost of a solved contactCSAT at or above your human baseline, cost below itCheap containment that hides low resolution and rising repeat calls

Each row deserves a note. The first row is about correct actions on money and policy. Test returns both inside and outside the window. A good agent honors the returns policy you set; it does not improvise one.

Catalog grounding is the row retail teams underweight. Ask the agent for a price, a stock count, and a return window. Every answer must trace to your catalog. A hallucinated stock level looks fluent and costs you a chargeback.

Tool-call correctness is where WISMO lives. The agent must call your order management system, pass the right order ID, and read the result back accurately. For the failure modes here, see our tool calling guide.

How to run a retail support voice agent vendor evaluation

Run the same process against every vendor. Identical inputs are the only way to compare fairly. Treat this like a structured request for proposal, scored on measured behavior rather than promises.

1. Define your scenarios from real calls. Pull your top intents from last quarter. Include order lookup, WISMO, returns, refunds, exchanges, and product questions. Add the messy ones: angry callers, wrong order numbers, expired promo codes.

2. Load your real catalog and policies. Give each vendor the same product data, stock feed, price list, and returns rules. Grounding cannot be tested on a demo catalog.

3. Connect the same tools. Wire each agent to test instances of your order, inventory, and loyalty systems. Tool calls are half of retail support.

4. Script the pass bars. Write the expected action for every scenario. Decide what "correct" means for a refund, an escalation, and a stock answer before you listen to a single call.

5. Run identical test suites. Send the same synthetic callers, in the same voices and accents, to every vendor. Vary phrasing so the agent cannot memorize.

6. Ramp to peak load. Push concurrency to your Black Friday peak and hold it. Watch latency, dropped calls, and accuracy as load climbs.

7. Score against the table. Mark each dimension pass, fail, or partial. Note every hallucination, every trapped caller, every wrong tool call.

8. Compare cost per resolution. Divide fully loaded cost by solved contacts, not by contained calls. Use the same denominator for every vendor.

9. Decide on the scorecard. Weight the dimensions by your priorities. Pick the vendor with the best measured behavior under load, not the best demo.

For a deeper build of the suites in steps 1 and 5, see our customer support testing guide.

Testing peak-season performance

Peak season voice agent testing is the dimension most demos skip. Vendors show you one call at a time. Your holiday traffic is thousands of calls at once. The right voice agent vendor for retail proves itself here, not in a scripted walkthrough.

Concurrency changes agent behavior. Under load, latency rises and models sometimes shed accuracy to keep up. A returns flow that worked at 200 calls can stall at 20,000. You need to see that curve before you sign.

Test it directly. Ramp synthetic callers from your normal volume to your peak, then hold the peak for at least an hour. Measure latency, dropped calls, and accuracy at each step. A vendor whose numbers hold within 15% of baseline is ready for December. One whose latency doubles is not.

Read peak testing as risk management. The NIST AI Risk Management Framework frames this well: measure a system under the conditions it will actually face. For retail, the condition that matters is your peak.

Grounding on your catalog, not the vendor's

Grounding is the retail-specific test that separates good vendors from risky ones. The agent must answer product, stock, price, and policy questions from your data alone.

Ask it hard questions. Request a price for an item on promotion. Ask whether a product is in stock in a specific size. Ask if an order placed 40 days ago can be returned. Every answer must match your catalog and your returns window.

Watch for confident invention. A weak agent fills gaps with plausible-sounding numbers. It quotes a stock level it never checked. It grants a return the policy forbids. Each of those is a hallucination, and each one costs money or trust.

Promo and loyalty handling belong here too. Test valid and expired coupon codes. Test loyalty program tier lookups. The agent should apply real rules, not guess at them.

Containment versus escalation

Containment is tempting to optimize and easy to fake. A high number looks like automation working. It can also mean callers are trapped.

The retail edge case is the angry caller. A customer whose order is lost during the holidays is not calm. The agent must stay steady, avoid over-promising, and escalate when it cannot help. A pushy upsell to that caller is a failure, not a feature.

Test escalation as carefully as resolution. Send calls the agent should not handle. It must hand off cleanly, with full context, to a human. An escalation that drops the caller's history forces them to repeat everything. For the patterns here, see our escalation and handoff guide.

The number to trust is true resolution, not raw containment. Resolution counts calls where the problem was actually solved. Containment counts calls that avoided a human, solved or not.

CSAT and cost per resolution

The last dimension pairs quality with cost. Neither number stands alone. Cheap calls that leave customers angry are not a win.

Measure customer satisfaction with post-call sentiment and short surveys. Compare each vendor to your human baseline. A voice agent should hold CSAT steady while it lowers cost, not trade one for the other.

Then measure cost per resolution, not cost per call. Use the total cost of ownership: platform fees, telephony, tool usage, and human escalations. Divide by solved contacts. That single number lets you compare vendors honestly.

Run these comparisons like an A/B test. Same scenarios, same catalog, same load, different vendor. Only the vendor changes, so differences in the results are real.

Where an independent evaluator fits

Every vendor grades its own homework. Their numbers come from their data, their scenarios, and their load. That is marketing, not evidence you can bet peak season on. So evaluate retail support voice agent vendor claims against your own catalog instead.

Evalgent is an independent, third-party evaluator for voice agents. We test each vendor on your catalog, your order flows, and your scenarios. We ramp concurrency to your Black Friday peak and hold it. We flag every hallucinated price, every trapped caller, and every wrong tool call.

The output is a neutral scorecard across all six dimensions. It tells you which vendor holds up under load and grounds on your catalog, and which one only shines in a demo. For the case behind neutral evaluation, see our independent voice AI evaluation explainer.

Book a demo to see a retail evaluation run on your own scenarios.

Frequently asked questions

How do you evaluate a retail support voice agent vendor?

Run identical tests against each vendor using your real catalog, order flows, and peak traffic. Score six dimensions: order and refund handling, catalog grounding, tool-call correctness, peak-load performance, containment and escalation, and CSAT with cost per resolution. Decide on measured behavior under load, not on a scripted vendor demo.

How do you test a voice agent for peak season?

Ramp synthetic callers from normal volume to your Black Friday peak, then hold that peak for at least an hour. Measure latency, dropped calls, and accuracy at each concurrency step. A vendor that stays within 15% of its low-load numbers is ready. One whose latency doubles under load is not.

How do you stop a voice agent from hallucinating stock or prices?

You cannot stop it from the vendor side, so you test for it. Ask price, stock, and returns-window questions answerable only from your catalog. Every answer must trace to your data. Any invented stock level, price, or policy is a hallucination and a red flag that should fail the vendor.

What should a retail voice agent do with an angry customer?

It should stay calm, avoid over-promising, and never push an upsell on a frustrated caller. When it cannot solve the problem, it must escalate cleanly to a human with full context. Test this directly with angry, edge-case calls. Trapping the caller inside the bot is a failure.

How do you measure cost per resolution for a voice agent?

Add the fully loaded costs: platform fees, telephony, tool usage, and human escalations. Divide that total by the number of contacts actually solved, not by contained calls. Use the same denominator for every vendor. Cost per resolution exposes cheap containment that hides low resolution and rising repeat contacts.

Can a voice agent handle order lookup and returns?

Good agents handle order lookup, returns, exchanges, and refund status by calling your order system with correct arguments. The risk is in edge cases: returns outside the window, wrong order numbers, and failed tool calls. Test those directly. A pass bar of 95% correct actions on scripted order scenarios is a reasonable start.

How do you test WISMO on a voice agent?

WISMO — "where is my order" — is a tool-call test. The agent must call your order management system, pass the correct order ID, and read back the real status. Test valid and invalid order numbers. A fabricated shipping status when the call fails is a red flag, not a graceful fallback.

What pass bar should a retail voice agent meet?

Reasonable starting bars: 95%+ correct actions on order scenarios, zero catalog hallucinations, 98%+ correct tool calls, latency within 15% of baseline at peak, clean escalation on 100% of out-of-scope calls, and CSAT at or above your human baseline. Set your own thresholds where your risk and margin differ.

The bottom line

A retail support voice agent vendor is only as good as it performs on your catalog, at your peak, on your hardest calls. Test all six dimensions on identical inputs, hold each vendor to the same pass bars, and decide on cost per resolution rather than a demo.

Related Articles