Test your voice agent
Metrics for an insurance claims voice agent

> Quick Answer: An insurance claims voice agent is measured on claim-data capture accuracy, FNOL completeness, compliance and disclosure accuracy, fraud-flag handling, correct-routing rate, escalation accuracy, and empathy. Completeness and accuracy on claim details drive downstream cost, so those anchors matter most before any convenience metric.
A claims voice agent lives or dies on the details it captures. A wrong policy number sends a claim to the wrong queue. A missing loss date stalls the file for days. A misheard address delays the adjuster visit. Every gap becomes rework, and rework is expensive.
So the metrics for a claims agent are not about how fast it talks. They are about whether the claim it opens is complete, accurate, compliant, and routed to the right place. This guide defines the KPIs, gives adaptable targets, and works through an example. For the how-to of building the test suite behind these numbers, see our insurance testing guide.
Why claim-data quality drives downstream cost
Every claims process starts with intake. The first notice of loss, or FNOL, sets the whole file in motion. If intake is clean, the rest flows. If intake is broken, the cost compounds.
A missing field triggers an outbound callback. That callback costs staff time and annoys the claimant. A wrong field routes the claim to the wrong adjuster. That misroute adds a handoff, a re-read, and a delay. Both failures started as one small data error at intake.
This is why data completeness and accuracy sit at the top of the scorecard. They are leading indicators. A convenience key performance indicator like handle time is a lagging one. A short call that captured half the claim is not a win. It is a callback waiting to happen.
The rule is simple. Measure the quality of the claim record first. Measure the speed of the call second. Speed still counts, since long gaps break the call. The ITU-T G.114 recommendation marks how much delay a caller notices. But optimizing speed before quality just produces fast, broken files. Our metrics hub covers the universal buckets that sit under this.
Claim-data capture accuracy
Capture accuracy measures whether the agent recorded each field correctly. Not whether it heard the caller, but whether the value written to the claim is right. These are different things.
Focus on critical entities. Policy number, claim type, loss date, loss location, and contact details drive routing and payment. An error in any of them breaks the file. Measure accuracy on these fields specifically, not as a blended average.
The relevant signal here is close to word error rate, but applied per field. An agent can post a low overall transcription error and still mangle the one policy number that matters. So score each critical field as pass or fail against ground truth.
A workable target is 99 percent or higher accuracy on critical entities. Names and street addresses are harder, so a slightly lower bar there is fair. But policy numbers and loss dates should be near perfect. Those two errors cause the most expensive downstream rework.
FNOL completeness
Completeness measures whether every required field was captured before the call ended. Capture accuracy asks if a field is right. Completeness asks if it is even there. A claim missing its loss description is incomplete, no matter how accurate the rest is.
Define the required set for each claim type. An auto claim needs the loss date, location, vehicles, injuries, and a description. A property claim needs the damage type, cause, and affected areas. Score the call on the fraction of required fields captured.
The target should be high, because gaps are costly. Aim for 95 percent or more of claims captured complete on the first call. Every incomplete claim is a callback, and callbacks are the biggest hidden cost in intake. Completeness is the metric that kills them.
Track completeness by claim type, not in aggregate. A blended number hides the type that is failing. If auto intake is clean but property intake drops fields, the average looks fine while property adjusters drown in callbacks.
Compliance and disclosure accuracy
Insurance intake carries legal duties. The agent may need to state recording notices, fraud warnings, or state-specific disclosures. It must not give advice it is not licensed to give. Compliance is a gate, not a nice-to-have.
Measure disclosure accuracy as the fraction of calls where every required statement was made, correctly and at the right time. A fraud warning read after the claim is logged may not satisfy the requirement. Timing counts, so score both presence and placement.
Also measure the unlicensed-advice rate. The agent should never speculate on coverage or promise a payout. It should state facts and defer judgment to a licensed adjuster. Any coverage opinion is a violation, and one violation can outweigh many clean calls.
Treat compliance as pass or fail with a hard threshold. A call below the disclosure bar blocks release, regardless of other scores. Frameworks like the NIST AI Risk Management Framework give useful structure for governing these controls.
Fraud-flag and suspicious-claim handling
Some claims carry warning signs. Inconsistent dates, a recent policy change, or a rehearsed narrative can signal a problem. The agent will not adjudicate insurance fraud. But it must capture the signals cleanly and flag them for review.
Measure two things here. First, the flag capture rate. Did the agent record the indicators the process defines, such as conflicting statements or unusual timing? Second, the false-flag rate. Did it flag ordinary claims and create needless review load?
The goal is high sensitivity without noise. A missed flag lets a suspicious claim slip through. A flood of false flags buries the special investigations team. Both cost money, in different ways. Tune the balance and track both numbers together.
Crucially, the agent must not accuse the caller. It records signals neutrally and routes for human review. Measure tone on flagged calls to confirm the agent stayed factual. An accusatory agent creates complaints and legal exposure, even when the flag was correct.
Resolution and correct-routing rate
For a claims agent, resolution rarely means closing the claim. It means completing intake correctly and sending the file to the right destination. So the honest outcome metric is correct-routing rate, not raw containment.
Measure the fraction of calls where the claim landed in the correct queue with a complete record. A claim routed to auto when it belongs in property is a misroute, even if intake was complete. A complete claim in the wrong place still costs a handoff and a delay.
For claim-status calls, resolution is cleaner. Did the caller get an accurate status without a callback? Measure the fraction answered correctly on the first contact. A caller who calls back the next day was not resolved, whatever the first transcript claimed.
Be wary of containment as a headline. Containment counts any call handled without a human. But a claim trapped and half-captured is not contained in any useful sense. Pair routing accuracy with callback rate to separate real self-service from dead ends. Our routing and intent guide goes deeper.
Escalation accuracy
Not every claim should stay with the agent. A total loss, a bodily injury, or a distressed caller needs a human. Escalation accuracy measures whether the agent handed off at the right moment, neither too early nor too late.
Score escalations against a defined policy. A hand-off that should have happened but did not is a miss. A hand-off that fired on a routine claim is a false escalation. Both hurt. Misses trap callers; false escalations waste adjuster time.
The target depends on risk tolerance. High-severity triggers, like injury or total loss, should escalate close to 100 percent of the time. Lower-stakes triggers can carry more slack. Set thresholds per trigger, not one blended rate.
Watch for the pattern where task success looks high but satisfaction is low. That gap usually means the agent pushed through when it should have transferred. Our escalation design guide covers how to define these triggers well.
Empathy and caller experience
Claimants often call on a bad day. A car crash, a flooded basement, or a theft is stressful. An agent that is accurate but cold still fails the moment. So empathy belongs on the scorecard, measured, not assumed.
Score empathy on acknowledgment, tone, and pacing. Did the agent acknowledge the loss before diving into fields? Did it avoid rushing a shaken caller? These are judgeable from the transcript and audio, and they correlate with downstream customer satisfaction.
Keep the measure concrete. Use a rubric with clear pass conditions, not a vague vibe score. For example, acknowledgment of the loss within the first two turns is a binary check. Consistent rubric scoring makes empathy a real metric, not an opinion.
Do not let empathy trade against accuracy. The best claims calls are warm and complete. A gentle agent that drops fields is not kind. It is a callback that felt nice for ninety seconds.
Adaptable targets at a glance
The table below sets a starting point. Treat these as defaults to tune against your own risk profile and claim mix, not fixed laws.
| Metric | Starting target | Why it drives cost |
|---|---|---|
| Critical-entity capture accuracy | 99%+ on policy number, loss date | Wrong values cause rework and misroutes |
| FNOL completeness | 95%+ complete on first call | Gaps become outbound callbacks |
| Disclosure accuracy | 100% required statements, correct timing | Missing disclosures create legal exposure |
| Unlicensed-advice rate | 0 coverage opinions | One opinion can trigger a complaint |
| Fraud-flag capture | High sensitivity, low false-flag noise | Misses and noise both add cost |
| Correct-routing rate | 95%+ to the right queue, complete | Misroutes add handoffs and delay |
| High-severity escalation | Near 100% on injury or total loss | Missed hand-offs trap vulnerable callers |
| Empathy rubric pass | Acknowledgment within two turns | Cold calls raise complaints and churn |
A worked example
Picture an auto FNOL agent handling a rear-end collision call. The caller is rattled. The claim has seven required fields and three critical entities. Here is how the metrics read on one call.
The agent acknowledges the crash in the first turn. Empathy passes. It reads the recording notice and fraud warning up front. Disclosure accuracy passes. It captures the policy number, loss date, and location correctly. Critical-entity accuracy is clean.
But it forgets to ask about injuries. That is a required field for auto claims. Completeness drops to six of seven, so this call fails the completeness bar. Because injury is a high-severity trigger, the missing question is also an escalation miss.
Now trace the cost. The missing injury field forces an outbound callback the next day. Staff time is spent. The claimant re-tells a stressful story. The adjuster assignment waits. One skipped question created three downstream costs.
This is why completeness and escalation sit above handle time. The call was fast, warm, and mostly accurate. It still generated rework. A scorecard that only tracked speed would have called this a success. The claim record tells the truth.
How to build a claims metrics scorecard
Follow these steps to turn the KPIs above into a working scorecard you can defend to operations and compliance.
1. List required fields per claim type. Write down every field a complete FNOL needs, split by auto, property, and other lines. This defines completeness and capture targets precisely.
2. Define ground truth for each field. For every test case, record the correct value. This lets you score capture accuracy as pass or fail against a known answer, not a guess.
3. Encode compliance as hard gates. Turn each disclosure and the unlicensed-advice rule into a binary check. Set a threshold below which a call blocks release, no exceptions.
4. Specify fraud and escalation triggers. Name the exact signals that should flag a claim or force a hand-off. Ambiguous triggers produce inconsistent scores and arguments later.
5. Write an empathy rubric. Define concrete pass conditions, like acknowledgment within two turns. Concrete beats vague, because two reviewers should reach the same score.
6. Segment every metric by claim type and caller profile. Report auto and property separately, and vary accents and noise. Blended averages hide the segment that is failing.
7. Run the scorecard on every change. Treat it as a regression suite, not a launch report. Re-run on each model swap, prompt edit, and integration update, so drift is caught before claimants feel it.
Comparing this to a production readiness bar
A metrics scorecard tells you how the agent scores today. A readiness bar tells you whether that score is good enough to ship. They are related but distinct, and both belong in an evaluation program.
The scorecard is a measurement. It produces numbers per metric, per segment. The readiness bar is a decision. It sets the minimum each number must clear before release, plus the hard gates that block a launch outright.
For claims, the gates are obvious. Disclosure accuracy and high-severity escalation are pass-or-fail. A strong average cannot buy back a compliance miss. Our production readiness guide explains how to set those thresholds, and our vendor evaluation guide covers comparing agents against a shared bar.
Frequently asked questions
What is the most important metric for a claims voice agent?
Claim-data completeness and critical-entity accuracy come first. A claim that is fast but incomplete becomes an outbound callback. Wrong policy numbers or loss dates cause misroutes and rework. These leading indicators drive downstream cost, so they outrank convenience metrics like handle time on any honest claims scorecard.
Why does FNOL completeness matter so much?
The first notice of loss sets the whole file in motion. A missing required field forces a callback, which costs staff time and re-traumatizes the claimant. Completeness is a leading indicator of downstream cost. Track it by claim type, since a blended average hides the specific line that is dropping fields.
How should a claims voice agent handle suspicious claims?
It should capture the defined fraud signals cleanly and flag the claim for human review, neutrally. It must never accuse the caller or adjudicate anything. Measure both flag capture rate and false-flag rate together. Missed flags let problems through; too many false flags bury the investigations team in noise.
Can a claims voice agent give coverage advice?
No. It should state facts and defer judgment to a licensed adjuster. Any coverage opinion or payout promise is a compliance violation. Measure the unlicensed-advice rate and treat it as a hard gate. One opinion can trigger a complaint, so a single violation outweighs many otherwise clean intake calls.
What is correct-routing rate for a claims agent?
It is the fraction of calls where the claim reached the right queue with a complete record. For a claims agent, this is a more honest outcome than raw containment. A claim routed to the wrong line still costs a handoff and delay, even when intake was captured completely and accurately.
How do you measure empathy on a claims call?
Use a concrete rubric, not a vague vibe score. Check binary conditions, like acknowledging the loss within the first two turns and avoiding rushing a shaken caller. These are judgeable from transcript and audio. Consistent rubric scoring makes empathy a real, defensible metric rather than a reviewer's opinion.
What escalation triggers should a claims agent have?
High-severity events like bodily injury or total loss should escalate close to every time. Distressed callers and ambiguous coverage questions also warrant a human. Set thresholds per trigger, not one blended rate. Score both misses and false escalations, since trapping a vulnerable caller and wasting adjuster time each carry cost.
How often should I run a claims metrics scorecard?
Run it as a regression suite on every change: model swaps, prompt edits, and integration updates. Segment results by claim type and caller profile each run. Drift in one segment hides inside a blended average, so a scorecard you run only at launch will miss the failure that reaches claimants first.
Measuring your claims agent with Evalgent
Evalgent is an independent platform for testing and evaluating voice agents against the metrics that actually matter. For a claims agent, it maps the KPIs above onto five primitives.
- Scenarios define each claim type and required field, so completeness and capture have a precise target.
- Profiles simulate rattled, accented, and noisy callers, so every metric segments the way real intake does.
- Metrics capture the anchors, including per-field accuracy, disclosure timing, routing, and the empathy rubric.
- Evaluations run the scorecard as a regression suite on every change, before any claimant is affected.
- Reviews let your team inspect the call behind each score, so a low number points to a real transcript.
Because Evalgent is independent, its scores are not tuned to flatter any single vendor. To see the scorecard run against your own claims agent, book a demo.
The bottom line
Measure claim-data completeness and accuracy first, because gaps and errors become the most expensive downstream rework. Speed and containment come second, always behind a complete, compliant, correctly routed claim.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more