Evalgent
Back to Blog
Voice AI Evaluation

Metrics for a collections voice agent

Deepesh Jayal
12 min read
Metrics for a collections voice agent

A collections voice agent lives inside one of the most regulated conversations in business. It has to recover money while staying inside the rules. That split defines how you measure it. Some metrics describe how well the agent collects, and others describe whether it stayed legal. The two are scored differently, and confusing them is how teams ship an agent that performs well and creates liability. This post lays out the KPIs, the targets, and which numbers are gates.

It pairs with our collections testing guide, which explains how to run the scenarios that produce these numbers. Here we focus on the numbers themselves.

Why collections metrics split into two classes

Most voice agent scorecards optimize a single axis: did the caller get what they wanted. Collections cannot work that way. A collections agent can hit a high resolution rate by pressuring people, disclosing debts to the wrong party, or skipping a required disclosure. Those shortcuts inflate the business number while breaking the law.

So the metrics divide into two classes. Business KPIs measure recovery and efficiency. You improve them gradually over time. Compliance KPIs measure legal conduct, and they are binary gates. A key performance indicator that trends upward is good news for recovery. But a compliance metric below 100% is a release blocker, not a trend to nudge. Keeping the two straight is the whole discipline.

The business KPIs that measure recovery

These are the numbers your collections leaders already watch, translated to an agent you can test.

Promise-to-pay rate is the share of eligible right-party conversations where the consumer commits to a payment or a plan. It is the closest proxy for the agent doing its core job. Resolution rate is broader. It is the share of calls that reach a defined, correct outcome — a payment, a plan, a documented dispute, or a clean stop request. Right-party-contact rate measures how often the agent reaches and verifies the correct consumer. A call to the wrong person can recover nothing and still risk a violation.

Alongside these, track handle time, responsiveness, and containment. Handle time shows efficiency. But a collections call should never be rushed at the cost of a disclosure. Responsiveness matters too, since long turn gaps read as evasive on a tense call. The ITU-T G.114 telephony standard puts the comfortable one-way limit at 150ms. Containment counts calls the agent finished without a human. But a "contained" call that pressured a distressed caller is a false win, not a success — the trap our containment versus deflection guide describes.

The compliance KPIs that gate a release

This is where collections differs from every other use case. These metrics are pass-or-fail, and the target is 100%.

Mandatory-disclosure accuracy measures whether the agent delivers required disclosures correctly, in the right place, on the right calls. In the United States that includes the mini-Miranda. That means identifying the call as an attempt to collect a debt, noting that information will be used for that purpose, and providing validation information. A missed disclosure is a violation. It is not a lower quality score. It fails the call outright.

One caution: this gate depends on the agent hearing the caller correctly. A high word error rate on noisy calls can cause a misheard name to pass verification or a stop request to be missed. So track transcription quality as an input to disclosure and right-party accuracy, not as a separate vanity number.

FDCPA compliance adherence is the broader gate. The Fair Debt Collection Practices Act and its implementing rules govern what a collector must say, must not say, and must not do. Testing turns each rule into an assertion: no threats, no harassment, no false statements, no debt disclosure to an unauthorized third party. The metric is the share of calls with zero violations. The acceptable rate is 100%.

Escalation accuracy is the third gate. A collections agent must hand off correctly when a caller disputes the debt, invokes a cease-and-desist, asks for a human, or hits a situation outside its authority. Measure whether it escalated when it should have, and — just as important — whether it kept collecting when it should have stopped. Getting this right is covered in our escalation guide.

Tone and empathy as a measured signal

Compliance sets the floor, but tone still matters. People in collections are often stressed, defensive, or hostile. An agent that stays professional under pressure protects the brand and reduces complaints. An agent that turns cold or combative invites both.

Score tone and empathy on a rubric across a range of caller profiles, from cooperative to abusive. This is a graded quality metric, not a gate, and it connects to broad customer satisfaction outcomes. But note the overlap with compliance: a harassing tone is not merely unpleasant, it can itself be a violation. Where tone crosses into prohibited conduct, the compliance gate takes over and the call fails.

A worked example: setting targets for a card-recovery agent

Consider an agent that calls consumers about past-due credit card balances. Here is how the metrics and targets come together in practice.

You define the business targets against your current baseline. Say human agents reach a promise-to-pay rate of 22% on right-party calls. You set the agent's release target at 20%, close to parity, and treat anything above as upside. Resolution rate targets 85% of calls reaching a correct, documented outcome. Right-party-contact rate targets your existing dialer performance, since the agent inherits the same contact data.

The compliance targets are not negotiable. Mandatory-disclosure accuracy must be 100% across the test suite. FDCPA adherence must show zero violations. Escalation accuracy must be 100% on dispute, cease-and-desist, and human-request scenarios. Tone scores a graded rubric, with a release floor you set — say, 4 out of 5 average, and no call below 3.

The gate logic is what makes this work. If disclosure accuracy is 99%, the agent does not ship. That holds even if promise-to-pay beats the human baseline. One missed disclosure in a hundred is one violation in a hundred calls. The business numbers earn the agent a look. The compliance numbers decide whether it launches. This is the same discipline described in our production-readiness bar.

Collections metrics and adaptable targets

Use this as a starting scorecard, then adjust the targets to your portfolio and jurisdiction.

MetricClassStarting targetHow to read it
Mandatory-disclosure accuracyCompliance gate100%Any miss is a violation, not a quality dip
FDCPA compliance adherenceCompliance gate100% (zero violations)Prohibited conduct fails the call outright
Escalation accuracyCompliance gate100%Correct handoff on disputes and stop requests
Right-party-contact rateBusiness KPIMatch dialer baselineReaching and verifying the correct consumer
Promise-to-pay rateBusiness KPINear human parityCore recovery outcome on right-party calls
Resolution rateBusiness KPI80–90%Calls reaching a correct, documented outcome
Tone / empathyGraded quality4/5 avg, none below 3Professionalism across caller profiles

The compliance rows never move off 100%. The business rows move with your baseline, your book, and your risk appetite.

How to build a collections metric scorecard

Follow these steps to turn the list above into a scorecard you can gate releases on.

1. Separate compliance from business first. Label every metric as a hard gate or a graded KPI before you set any target, so the two classes never get averaged together into one score.

2. Set compliance gates at 100%. Disclosure accuracy, FDCPA adherence, and escalation accuracy each pass only with zero failures across the suite. Do not soften these into percentages you can trend.

3. Anchor business targets to your baseline. Pull current human promise-to-pay, resolution, and right-party rates, then set the agent's release target at or near parity rather than an arbitrary number.

4. Define each outcome precisely. Write down what counts as a resolution, a promise-to-pay, and a verified right party, so every call is scored the same way every time.

5. Score tone across caller profiles. Rate professionalism on a rubric spanning cooperative to hostile callers, with a release floor and a hard minimum for any single call.

6. Build the gate logic. Encode that a compliance failure blocks release regardless of business performance, so a strong recovery number can never buy back a violation.

7. Re-measure on every change. Run the full suite before each release, since a prompt tweak that lifts promise-to-pay can quietly break a disclosure. Compare against the metrics scorecard format for consistency.

Common metric mistakes in collections

The most frequent error is treating disclosure and conduct as quality scores. Teams report "97% disclosure compliance" as if it were a good grade. In collections, that reads as three violations per hundred calls, and it should block the release. Compliance is binary by nature, and averaging it hides the failures that create legal exposure.

A second mistake is optimizing promise-to-pay in isolation. An agent tuned only to secure commitments will learn to pressure, and pressure invites both complaints and violations. Recovery numbers must always be read next to the compliance gates and the tone score, never alone.

A third is scoring only the happy path. Real collections calls include disputes, wrong numbers, hostile callers, and stop requests. Metrics measured only on cooperative calls describe an agent that does not exist once real traffic arrives. When you compare vendors on these calls, use the same discipline as our vendor evaluation guide, scoring every agent on identical scenarios.

Aligning metrics to a governance framework

Regulated collections teams increasingly need to show that their measurement maps to a recognized standard, not just internal preference. Aligning your metric definitions to the NIST AI Risk Management Framework gives compliance and legal a shared vocabulary for how the agent's risks are measured and gated. That matters when a call is later questioned.

Records are part of the metric picture too. Collections is audited, so each call needs a complete trail: what was disclosed, who was verified, and what the consumer requested. Confirm the agent produces that record and that sensitive details in it are handled correctly. A metric you cannot evidence later is a metric a regulator will not accept.

Measuring a collections voice agent with Evalgent

Evalgent measures collections agents where the risk actually lives — on hard, adversarial calls, with compliance and business KPIs scored separately. Scenarios script right-party verification, required disclosures, disputes, cease-and-desist requests, and hostile callers, so the compliance moments are exercised rather than assumed. Profiles vary caller tone from cooperative to abusive, since composure under pressure is part of compliance. Metrics encode disclosure accuracy, FDCPA adherence, and escalation accuracy as pass-or-fail gates, with promise-to-pay and resolution as graded KPIs you target against your baseline. Evaluations run the full suite as automated batches before every release. Reviews let your compliance team replay any call with audio, transcript, and metrics together, which is where a collections agent gets signed off.

The result is a scorecard your compliance, legal, and collections teams can all approve: recovery measured honestly, disclosures made every time, and violations gated to zero. To measure your collections agent against these gates, book a demo.

The bottom line

Collections voice agent metrics come in two classes: business KPIs you improve over time, and compliance KPIs that pass or fail at 100%. Score recovery honestly, but never let a strong promise-to-pay number buy back a missed disclosure — that is a violation, not a quality ding.

Frequently asked questions

What metrics matter most for a collections voice agent?

Two classes matter. Business KPIs — promise-to-pay rate, resolution rate, and right-party-contact rate — measure recovery. Compliance KPIs — mandatory-disclosure accuracy, FDCPA adherence, and escalation accuracy — measure legal conduct. The compliance metrics are hard pass-or-fail gates set at 100%, while the business metrics are targets you improve against your baseline over time.

Why is disclosure accuracy a pass-or-fail metric?

Because a missed mandatory disclosure is a legal violation, not a lower quality score. In the United States the mini-Miranda and validation disclosures are required by law, so an agent that skips one on a call has broken a rule on that call. Averaging it into a percentage hides the violations, so the target is 100%, and any miss blocks release.

What is a good promise-to-pay rate for a voice agent?

Anchor it to your human baseline rather than an absolute number. If your collectors reach commitments on roughly a fifth of right-party calls, set the agent's release target near parity and treat anything higher as upside. The right target depends on your portfolio, your consumers, and your risk appetite, so measure your own baseline before setting one.

How do you measure right-party-contact rate?

Track the share of calls where the agent reaches and correctly verifies the intended consumer before discussing any debt detail. Score both sides: correct verification of the right party, and correct withholding when the wrong party answers. It is a business KPI for recovery, but it overlaps with compliance, since disclosing a debt to the wrong person can itself be a violation.

How do you measure escalation accuracy in collections?

Run scenarios where a caller disputes the debt, invokes a cease-and-desist, asks for a human, or hits a limit of the agent's authority. Assert the agent escalates or stops when it should, and keeps handling when it should. Escalation accuracy is a gate, targeted at 100%, because failing to hand off a dispute or stop request is a compliance failure.

Is tone a compliance metric or a quality metric?

Both, depending on severity. Tone and empathy are normally graded quality metrics, scored on a rubric across cooperative to hostile callers. But when tone crosses into harassment or abusive conduct, it becomes a compliance failure and the gate takes over. So score tone as a graded number with a release floor, while treating prohibited conduct as a hard, call-failing violation.

Can a collections agent ship if it beats the recovery baseline?

Not if it fails a compliance gate. Strong promise-to-pay and resolution numbers earn the agent a serious look, but they cannot buy back a missed disclosure or a mishandled cease-and-desist. The gate logic is deliberate: business KPIs decide whether the agent is good, and compliance KPIs decide whether it is allowed to launch at all.

How often should collections metrics be re-measured?

Before every release, and after any change to the prompt, model, or flow. A tweak that lifts promise-to-pay can quietly break a disclosure or an escalation path, so recovery gains never justify skipping the compliance re-run. Run the full suite each time, since collections metrics are only meaningful when the compliance gates are re-verified alongside the business numbers.

Related Articles