Evalgent
Back to Blog
Voice AI Evaluation

What's the North-Star Metric for a Voice Agent?

Deepesh Jayal
12 min read
What's the North-Star Metric for a Voice Agent?

# What's the North-Star Metric for a Voice Agent?

Quick answer

A voice agent's north star metric is verified task success: the share of calls where the caller's actual goal got done. It beats vanity metrics like containment or deflection, which can rise while service degrades. Pair it with guardrails for quality, safety, and latency, and measure it independently.

Most voice agent teams track a dozen numbers and steer by none of them. The dashboard fills with containment, average handle time, and sentiment. None of these tells you whether callers left happy. A team without a single guiding metric optimizes everything a little and nothing well.

A north star metric fixes that. It names the one outcome the whole team moves toward. For a voice agent, that outcome is almost always verified task success. This post explains what makes a good north star, why vanity metrics fail the job, which guardrails protect the number, and why the north star must be measured by an independent party.

North star metric: the single number a team optimizes toward because it best captures the value delivered to the customer and the business. Everything else is a supporting or guardrail metric.

What a north star metric is and is not

A north star metric is one number that predicts durable value. It sits above the day-to-day dashboard. Teams borrow the idea from product analytics and key performance indicators. The point is focus. One metric aligns engineering, operations, and leadership.

The concept pairs naturally with goal-setting frameworks like OKRs. The north star is the objective. The other metrics are how you get there. Without a north star, teams chase whichever number moved last week.

A north star is not the only metric you track. It is the metric you optimize for when two numbers conflict. When latency improves but task success drops, the north star tells you to stop. That trade-off clarity is the whole reason to pick one.

It is also not a vanity number. A vanity metric looks good in a slide but does not predict value. Call volume handled is a vanity metric. It rises whether callers were helped or not. A real north star cannot move up while customers suffer.

Why verified task success is the north star

Verified task success is the share of calls where the caller's goal was completed. The refund was issued. The appointment was booked. The claim was filed. You confirm the outcome against evidence, not the agent's own claim.

This metric ties directly to the caller. Callers do not phone in to be contained. They phone in to get something done. When task success rises, callers are leaving with their problem solved. That is the value the business is paying for.

Academic work on dialogue systems reached this conclusion decades ago. The PARADISE framework anchored dialogue quality to task completion, as Walker and colleagues described in 1997. The insight predates modern voice agents. Completion is still the outcome that matters most.

Task success also resists gaming better than the alternatives. An agent cannot inflate it by stalling, deflecting, or sounding confident. The goal either got done or it did not. Our task success rate guide covers how to define and label it per call type. Our metrics scorecard shows how it sits at the center of the wider metric set.

Vanity metrics that fail as the single goal

Containment is the most common false north star. It measures the share of calls the agent handled without a human transfer. Teams like it because it looks like automation working. It is a trap when used as the top-line goal.

Containment rises when the agent refuses to transfer. It rises when callers give up. It rises when the agent talks in circles until the caller hangs up. None of those outcomes helped anyone. A contained call can still be a total failure. Our containment rate guide breaks down where the number misleads.

Deflection has the same flaw. It counts calls pushed away from live agents. Push a caller to a dead-end web form and deflection improves. The caller's problem stays unsolved. The containment versus deflection guide separates these two terms, which teams mix up constantly.

This is Goodhart's law in action. When a measure becomes a target, it stops being a good measure. Make containment the goal and the agent learns to hold callers hostage. The number climbs while service collapses. That is exactly the failure a north star is supposed to prevent.

The deeper problem is that containment is a proxy) for success, not success itself. Proxies drift from the thing they stand in for. When you optimize the proxy directly, the gap widens. Task success avoids this because it measures the outcome, not a stand-in for it.

What makes a good north star metric

A good north star metric passes three tests. It ties to a real caller or business outcome. It is hard to game. It is measurable on real traffic. Weak candidates fail at least one of these.

The table below scores common candidates against the three tests. It shows why verified task success wins and why the popular alternatives fall short as the single goal. Use it to sanity-check any metric someone proposes as your north star.

Candidate metricTied to caller outcome?Hard to game?Verdict as the single goal
Verified task successYes, measures goal completionYes, outcome is checked against evidenceBest north star for most teams
ContainmentNo, only measures no-transferNo, rises when callers give upFails; use as an efficiency guardrail
DeflectionNo, only measures channel shiftNo, rises with dead-end routingFails; use as a channel metric
Average handle timeWeakly, speed is not successNo, rewards rushing callers offFails; use as an efficiency guardrail
Call volume handledNo, activity not outcomeNo, rises regardless of qualityVanity metric; never a north star

The pattern is clear. Metrics that measure activity or efficiency make bad goals. They reward the agent for doing something, not for helping. Only an outcome metric survives all three tests. That is why verified task success is the default north star for voice agents.

Pick your north star for the outcome your callers came for. A support line optimizes for issue resolution. A booking line optimizes for confirmed appointments. A collections line optimizes for completed arrangements. The shape of "success" changes by use case, but it is always the caller's goal.

Guardrail metrics that protect the number

A north star alone is dangerous. Optimize any single metric hard enough and you can degrade everything around it. Guardrail metrics catch that. They are the numbers that must not get worse while you push the north star up.

Think of guardrails as constraints. The team maximizes task success subject to guardrails staying healthy. If a change lifts success but breaks a guardrail, it does not ship. This keeps you from winning the number and losing the customer.

Guardrail metric: a supporting number that must stay within a healthy range while you optimize the north star. It catches the damage a single-metric focus can hide.

Four guardrails matter most for voice agents.

Customer satisfaction, or CSAT, catches quality damage. An agent could force task success up by being pushy or robotic. CSAT flags when the experience got worse even as goals got done. It is the caller's own verdict on the interaction.

Escalation accuracy catches the opposite failure. When the agent cannot help, it should hand off cleanly and early. This guardrail measures whether hard calls reach a human at the right moment. It stops the agent from trapping callers to protect its own numbers.

Latency catches the experience tax. Long silences and slow responses frustrate callers even on successful calls. Track response time and dead air. A win that takes ninety seconds of awkward pauses is a fragile win.

Safety and compliance catch the expensive failures. In regulated calls, the agent must follow policy, disclose what it must, and avoid banned claims. One compliant-looking success that broke a rule can cost more than a hundred clean wins. Treat policy adherence as a hard guardrail, not a nice-to-have.

Why the north star must be measured independently

A north star you grade yourself is a north star you will flatter. This is the core problem with self-reported metrics. The party that built the agent has every incentive to score its successes generously.

Self-grading fails in predictable ways. Success criteria get loosened until borderline calls pass. The agent's own "I've resolved that for you" gets counted as proof. Failed calls get filtered out as edge cases. None of this is malicious. It is the natural pull of measuring your own work.

Independent measurement removes that pull. A third party labels calls against fixed criteria, from the transcript, the audio, and your system of record. The score does not bend to protect anyone's roadmap. Our independent evaluation guide explains why separation of grader and builder matters.

Independence also makes the number comparable over time and across vendors. When the same neutral rubric grades every wave, the north star means the same thing in March and September. You can benchmark two agents on the same test cases, as our own-data benchmarking guide describes. Evalgent is the independent, third-party evaluator that measures verified task success and your guardrails on your real calls.

There is a difference between testing and evaluation here. Testing checks whether the agent works before you ship. Evaluation measures whether it delivers outcomes in production. The testing versus evaluation guide draws the line. Your north star lives on the evaluation side, measured on live traffic.

How to choose your north star metric and guardrails

Follow these steps to set a north star your team can actually steer by.

1. Name the caller's goal for each call type. Write it as a completed outcome. "Refund issued," not "refund discussed." This becomes your definition of success.

2. Pick verified task success as the default north star. Define it as the share of calls where that goal got done, confirmed against evidence rather than the agent's claim.

3. Score candidates against the three tests. Confirm your chosen metric ties to a caller outcome, resists gaming, and is measurable on real traffic. Drop any candidate that fails one.

4. Choose three to five guardrails. Include CSAT, escalation accuracy, latency, and, in regulated cases, policy adherence. These are the numbers that must not get worse.

5. Set thresholds for each guardrail. Decide the range each must stay within. A change that breaks a guardrail does not ship, even if task success rose.

6. Assign independent measurement. Have a third party label calls against fixed criteria so the score cannot bend to protect anyone's roadmap.

7. Review the trade-offs on a fixed cadence. Each period, check the north star against guardrails. Investigate any gain that came at a guardrail's expense.

Start narrow. Pick one call type, define success cleanly, and measure it well before expanding. A precise north star on one flow beats a vague one across all of them.

Frequently asked questions

What is the north star metric for a voice agent?

The north star metric for a voice agent is verified task success. It is the share of calls where the caller's actual goal was completed, confirmed against evidence rather than the agent's own claim. It captures the value callers came for and resists gaming better than efficiency metrics like containment or deflection.

How to choose a north star metric for a voice agent?

Name the caller's goal for each call type as a completed outcome. Default to verified task success. Confirm it ties to a caller outcome, resists gaming, and is measurable on real traffic. Then add three to five guardrails, such as CSAT and latency, that must not degrade while you optimize.

Is containment a good north star metric?

No. Containment measures the share of calls handled without a human transfer, not whether callers were helped. It rises when callers give up or get stuck, so optimizing it directly can trap callers while service degrades. Use containment as an efficiency guardrail, and make verified task success the north star instead.

Why is task success rate the north star for voice agents?

Task success rate measures the outcome callers actually want: their goal getting done. It ties directly to customer and business value. It resists gaming because the goal either completed or it did not. Dialogue research has anchored quality to task completion for decades. That makes it the strongest single metric for a voice agent.

What guardrail metrics protect a voice agent north star?

Four guardrails matter most: customer satisfaction to catch quality damage, escalation accuracy to ensure clean handoffs, latency to catch slow or awkward experiences, and policy adherence for regulated calls. Guardrails are constraints. The team optimizes task success while keeping each guardrail within a healthy, predefined range.

Should you optimize for containment or task success?

Optimize for task success and treat containment as a guardrail. Containment measures efficiency, not outcomes, and can rise while callers leave unhelped. An agent that contains every call but solves nothing is worse than one that transfers often and resolves problems. Track both numbers, but steer by task success.

Can a north star metric be gamed?

Any metric can be gamed once it becomes a target, an effect known as Goodhart's law. Proxy metrics like containment are especially easy to inflate. Outcome metrics like verified task success are harder to game because the result is checked against evidence. Independent measurement and guardrails reduce the remaining gaming risk.

Who should measure a voice agent's north star metric?

An independent third party should measure it. The team that built the agent has an incentive to grade its own successes generously, loosening criteria over time. A neutral evaluator labels calls against fixed rubrics from the transcript, audio, and system of record, so the score stays comparable across waves and vendors.

The bottom line

The north star metric for a voice agent is verified task success, the share of calls where the caller's goal actually got done. Wrap it in guardrails like CSAT, escalation accuracy, and latency, and have an independent party measure it so the number cannot be gamed.

Pick the one outcome your callers came for and optimize toward it, not toward vanity metrics that flatter the dashboard. Book a demo to see how Evalgent measures verified task success and your guardrails on your real calls.

Related Articles