Test your voice agent
How to Measure First Contact Resolution (FCR)

# How to measure first contact resolution (FCR) for voice agents
Quick answer
First contact resolution (FCR) for a voice agent is the share of issues resolved on the first contact, with no callback, transfer, or repeat within a set window. Measure it on real calls, track repeats across every channel, and confirm the caller's problem was actually solved, not just ended.
Most teams report a resolution rate and stop there. That number is easy to inflate. A call can end cleanly and still fail the caller, who then calls back an hour later. First contact resolution is the metric that catches this. It asks a harder question: did the issue actually get solved the first time?
This guide explains what FCR means for a voice agent. It covers how to pick a measurement window, how to detect repeat contacts across channels, and why FCR exposes fake resolutions that a raw resolution rate hides. Evalgent measures FCR on your real calls as an independent third party, so the number reflects outcomes rather than optimism.
What first contact resolution means for a voice agent
> First contact resolution (FCR): the percentage of caller issues fully resolved on the first interaction, with no follow-up contact needed on the same issue. It measures whether the problem was solved, not whether the call ended.
The idea comes from the contact center world. First-call resolution has long been treated as a core quality signal in customer service. A high FCR means callers get what they need in one try. A low FCR means people call back, get transferred, or give up.
For a voice agent, the definition needs to be sharp. Three things break FCR, and all three must be absent for a contact to count as resolved:
- No callback. The caller does not contact you again about the same issue.
- No transfer. The agent does not hand the call to a human to finish the job.
- No repeat. The same intent does not resurface on another channel soon after.
Notice what this excludes. An agent that ends every call politely can still have terrible FCR. Ending a call is not the same as solving a problem. That gap is where most voice agent metrics go wrong.
FCR is a lagging outcome, not a call event. You cannot know if a contact was resolved at the moment it ends. You only know later, once the window closes and no repeat appears. This is why FCR is harder to measure than talk time or containment, and why it is worth the effort.
Why FCR exposes fake resolutions
A resolution rate counts calls the agent marked as done. FCR counts problems that stayed done. The difference is everything.
Consider a caller asking to change a shipping address. The agent confirms the change and ends the call. Resolution rate: one for one. But the change never wrote to the order system. The caller notices, calls back, and now a human fixes it. The real outcome was a failure that a resolution rate recorded as a win.
This is a fake resolution. The call looks successful in isolation. Only the repeat contact reveals the truth. A vendor reporting resolution rate alone will never surface it. An FCR measurement built on repeat detection will.
Fake resolutions cluster around a few patterns. The agent gives a plausible-sounding answer that is wrong. It confirms an action that did not complete. It ends the call to avoid a hard question. Each one inflates a naive success metric and quietly erodes trust. Poor resolution drives repeat business losses, because a caller who has to call twice trusts you less each time.
This is also why containment can mislead. A contained call is one a human never touched. But a contained call that failed the caller is worse than an escalation, because the failure is hidden. Our containment vs deflection guide unpacks how these terms get blurred. FCR cuts through the blur by tracking what happened next.
Choosing the measurement window (24h vs 7d)
FCR needs a window. Without one, "resolved" has no meaning, because any call could get a follow-up eventually. The window defines how long you wait before crediting a contact as resolved.
Two windows are common. A 24-hour window credits a contact as resolved if no repeat arrives within one day. A 7-day window waits a full week. Each has trade-offs.
A short window flatters the number. Many repeat contacts arrive on day two or day three, once the caller realizes the problem was not fixed. A 24-hour window misses those and reports FCR higher than reality. A 7-day window catches more genuine repeats, so the number is lower but more honest.
Pick the window that matches your issue types. Simple, same-day intents like checking a balance suit a 24-hour window. Issues that surface later, like a billing correction or a delivery that never arrives, need 7 days. When in doubt, measure both and report both. Hiding behind the shorter window is a form of gaming.
Whatever window you choose, apply it consistently. Changing the window between reports makes trend lines meaningless. Treat FCR as a performance indicator with a fixed, documented definition, the same way a mature contact center would.
Detecting repeat contacts across channels
FCR lives or dies on repeat detection. If you cannot tell that a caller came back, you cannot measure it. This is harder than it sounds, and it is where most in-house measurement quietly fails.
The first challenge is identity. A repeat only counts if you can link the second contact to the first. That means matching on a stable identifier: a phone number, an account ID, or a customer record. Free-text matching alone is unreliable. Build the link on identity, then confirm the intent matches.
The second challenge is intent. A caller who resolves a billing issue and later calls about shipping is not a repeat. Two different problems, two separate contacts. A true repeat is the same intent resurfacing. So repeat detection needs both a matched identity and a matched issue. Scoring the conversation intent is part of this, which our guide on how to score a voice agent conversation covers in depth.
Cross-channel leakage
The hardest challenge is cross-channel leakage. A caller talks to your voice agent, gets a fake resolution, and then opens a chat or sends an email instead of calling back. If you only watch the phone channel, that repeat is invisible. Your voice FCR looks great while the problem leaked to another queue.
This is a common and expensive blind spot. Voice agents that measure FCR on phone-only data systematically overstate it. The repeats did not vanish. They moved. Real FCR measurement joins contacts across voice, chat, email, and any other channel a caller can use.
To catch leakage, unify contacts at the customer level, not the channel level. Ask a simple question of every voice contact: did this customer reach out again on any channel about this issue within the window? If yes, it was not a first contact resolution, no matter how the call sounded.
FCR pitfalls and how to measure around them
FCR has a handful of well-known traps. Each one makes the number look better than the truth. The table below maps each pitfall to how it fools you and how to measure correctly.
| Pitfall | How it fools you | How to measure it correctly |
|---|---|---|
| Short window | A 24-hour window misses day-two and day-three callbacks, so FCR reads high. | Measure over 7 days for issues that surface late; report both windows side by side. |
| Channel switch | A caller who fails on voice reopens via chat or email, so voice-only FCR looks clean. | Join contacts at the customer level across every channel, not per channel. |
| Reopened ticket | A ticket marked resolved is reopened later, but the first contact still counts as a win. | Treat any reopen on the same issue within the window as a failed FCR. |
| Transfer | An agent hands off to a human, and the handoff gets logged as a resolution. | Count only contacts the agent finished alone; a transfer is never an FCR. |
Each of these pitfalls shares a root cause. The measurement stops too early or looks too narrowly. Fixing them means widening the lens in time and across channels, then holding the definition steady.
Transfers deserve a note. A clean escalation is good behavior, not a failure of judgment. The agent recognizing its limit and handing off is often the right call. But it is not a first contact resolution, and it should not be counted as one. Our escalation guide explains how to test that handoffs fire on the right triggers.
How to measure first contact resolution step by step
Here is a repeatable process for measuring FCR on a voice agent. It works whether you run the measurement yourself or bring in an independent evaluator.
1. Define resolved. Write down exactly what counts as resolved for each intent. The caller's stated goal must be met, with no callback, transfer, or repeat.
2. Set the window. Choose 24 hours or 7 days per issue type. Document it and do not change it between reports.
3. Capture real contacts. Pull actual production calls, not demo scripts. Include the messy ones: accents, interruptions, and edge cases.
4. Link identity. Match every contact to a stable customer identifier so repeats can be found.
5. Score the intent. Label each contact's intent so a later contact can be checked against it for a true match.
6. Join across channels. Merge voice, chat, and email at the customer level to catch cross-channel leakage.
7. Detect repeats. For each resolved contact, check the window for a same-intent repeat on any channel.
8. Exclude transfers. Remove any contact the agent handed to a human from the resolved count.
9. Compute FCR. Divide clean first-contact resolutions by total eligible contacts. Report the window and the sample size.
10. Segment and review. Break FCR down by intent, and read the failed cases to find fake resolutions.
The last step matters most. The FCR number tells you where to look. Reading the failed contacts tells you why they failed. That is where improvement comes from.
Where an independent evaluator fits
Measuring your own FCR has a conflict of interest. The team that built the agent picks the window, defines resolved, and decides which contacts count. Every one of those choices can nudge the number up. That is not always deliberate. It is just what happens when the scorer has a stake in the score.
Evalgent measures FCR as an independent third party. We define resolved against the caller's actual goal, apply a fixed window, and detect repeats across every channel using your real data. Because we do not build the agent, we have no reason to flatter it. The number reflects outcomes, not incentives.
This matters most when you compare vendors or track progress over time. An independent measurement is one both sides can trust. Our explainer on independent voice AI evaluation covers why third-party measurement changes the conversation. If you are assembling a broader metric set, the voice agent metrics scorecard shows how FCR fits alongside containment, escalation, and quality scores.
FCR is one metric in a system. It pairs well with a full voice agent evaluation that grades task success, safety, and conversation quality. For support-specific teams, our guide to customer support voice agent metrics puts FCR in context with the numbers that predict customer satisfaction. And when you want proof on your own traffic, running the measurement on your data, as described in benchmarking voice agents on your own data, is the only version that counts.
Task success measurement has a solid research base. Work on task-oriented dialogue systems, such as the widely cited MultiWOZ dataset, formalized how to judge whether a conversation actually completed the user's goal. FCR applies the same discipline to production calls.
Frequently asked questions
What is first contact resolution for a voice agent?
First contact resolution for a voice agent is the percentage of caller issues fully solved on the first interaction. No callback, transfer, or repeat happens on the same issue within a set window. It measures whether the problem was resolved, not whether the call ended politely.
How is FCR different from resolution rate?
Resolution rate counts calls the agent marks as done. FCR counts problems that stayed done, checked against later contacts. Resolution rate is easy to inflate with fake resolutions. FCR catches them by tracking repeats across a window and across channels, so it reflects real outcomes rather than call completion.
What measurement window should I use for FCR?
Use 24 hours for simple, same-day issues and 7 days for issues that surface later, like billing or delivery. A short window flatters the number by missing day-two callbacks. When unsure, measure both windows and report both. Keep the window fixed between reports so trends stay comparable.
Why does FCR expose fake resolutions?
A fake resolution is a call that looks successful but did not solve the caller's problem. The agent gives a wrong answer or confirms an action that never completed. Resolution rate records it as a win. FCR catches it because the caller comes back, and that repeat contact fails the first-contact test.
How do you detect repeat contacts across channels?
Link every contact to a stable customer identifier, then match on intent. Join voice, chat, and email at the customer level, not per channel. Check whether the same customer raised the same issue again within the window on any channel. Phone-only measurement misses repeats that leak to chat or email.
Does a transfer to a human count as first contact resolution?
No. A transfer means the agent did not finish the job alone, so it cannot count as a first contact resolution. A clean escalation is still good behavior when the issue exceeds the agent's scope. Exclude transferred contacts from the resolved count rather than logging the handoff as a success.
What is cross-channel leakage in FCR measurement?
Cross-channel leakage happens when a caller fails on voice, then reopens the issue via chat or email instead of calling back. Phone-only FCR misses these repeats and reads too high. The problem did not vanish; it moved queues. Unified, customer-level measurement across every channel is the only way to catch it.
Can I measure FCR on my own, or do I need an independent evaluator?
You can measure FCR in-house, but the team that built the agent also picks the window and defines resolved. Those choices can quietly inflate the number. An independent evaluator applies a fixed definition on your real data with no stake in the result, producing a figure both buyers and vendors can trust.
The bottom line
First contact resolution measures whether a voice agent solved the problem, not whether the call ended. Measure it on real calls, over a fixed window, across every channel, with transfers excluded.
Ready to see the FCR you can defend to buyers and your board? Book a demo and Evalgent will measure it independently on your real calls.
Related Articles

Why AI voice agents fail in production (and how to prevent it)
AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means.
Read more
Voice agent regression testing: why LLM updates break production
LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change.
Read more