Evalgent
Back to Blog
Voice AI Evaluation

How to Measure Mean Opinion Score (MOS) for Voice

Deepesh Jayal
12 min read
How to Measure Mean Opinion Score (MOS) for Voice

# How to measure mean opinion score (MOS) for voice agents

Quick answer

Mean opinion score (MOS) rates the naturalness of a voice agent's speech on a 1 to 5 scale. It comes from human listeners rating short audio clips, following the ITU-T P.800 method. A trustworthy MOS uses your own scripts, enough raters, randomized clips, and phone-quality audio, not a vendor's studio demo.

Most teams first meet MOS as a single number in a vendor deck. "4.5 MOS," it says, next to a waveform. It looks like proof the voice sounds human. It is not proof of anything until you know how it was measured.

MOS is a subjective score. It measures a feeling, not a fact. That makes it powerful and easy to abuse in equal measure. This piece shows how to measure a real mean opinion score for voice agents, how to run a proper listening test, and why a number from one study cannot be compared to a number from another.

It builds on our wider work on text-to-speech evaluation for voice agents, and on where a naturalness metric sits on a full voice agent metrics scorecard.

What mean opinion score actually measures

Mean opinion score measures perceived quality. For a voice agent, that quality is usually naturalness: does the synthetic voice sound like a human, or like a machine reading text? The score is the mean of ratings from many listeners.

> Mean opinion score (MOS): the arithmetic mean of subjective quality ratings, each given by a human listener on a 1 to 5 scale. It summarizes how good a group of people judged a set of audio samples to be.

The metric comes from telephony. Engineers needed a way to compare call quality across codecs and networks. The mean opinion score was the answer: play a clip, ask people to rate it, average the ratings. The same idea now rates modern speech synthesis, the technology behind a voice agent's spoken replies.

The scale is fixed and simple. Five means excellent. Four means good. Three means fair. Two means poor. One means bad. A listener picks one number per clip. You average across listeners and clips to get the MOS.

The key thing to hold onto is this. MOS is an opinion, measured carefully. It is not a physical measurement of the audio. Two studies can play the same clip and report different scores, because the people and the setup differed.

The 1 to 5 scale and absolute category rating

The standard way to collect MOS is called absolute category rating. Listeners hear one clip at a time. They rate it on its own, against the five-point scale, with no reference clip to compare against.

This makes each rating a judgment of the clip in isolation. The listener is not asked "which is better, A or B." They are asked "how good is this one." That absolute framing is what the ITU-T P.800 recommendation defines for telephony speech quality tests. You can read the scope of ITU-T P.800 for the formal method.

The five-point scale is a Likert scale. That has a consequence people forget. The steps are ordered, but they are not guaranteed to be equal. The gap between a 4 and a 5 may not feel the same as the gap between a 2 and a 3. So a MOS of 4.0 is not "twice as good" as a MOS of 2.0.

Absolute category rating is the common choice, but not the only one. Some tests use a comparison method, where listeners hear two clips and pick the better one, or rate the difference. Comparison tests are more sensitive when two voices are close. They answer a different question, though. They tell you which is better, not how good either one is on the standalone scale.

Why vendor MOS numbers are not comparable

Here is the single most important fact about MOS. A score is only meaningful inside the study that produced it. You cannot line up two vendor numbers from two different studies and call the higher one better.

The reason is that MOS has no fixed anchor. The score floats with the setup. Change the listeners, the clips, the instructions, or the playback device, and the same voice earns a different number. Nothing in the number itself tells you which setup produced it.

Consider what varies between studies. One vendor may use twenty expert listeners in a quiet room with studio headphones. Another may use two hundred crowd-workers on laptop speakers. One may read curated marketing sentences. Another may read hard cases full of numbers and names. Same 1 to 5 scale, completely different meaning.

Even the reference set shifts the number. If a study includes a very bad anchor clip, listeners rate everything else higher by contrast. Drop that anchor and the scores compress. This is why a "4.5 MOS" with no method attached is closer to marketing than measurement.

The fix is not to distrust MOS. It is to demand the method. A comparable MOS is one you ran yourself, on one setup, across all the voices you are judging. That is the only way the numbers sit on the same ruler. This is the same principle we cover in benchmarking voice agents on your own data.

MOS pitfalls and how to do it right

Most bad MOS numbers come from a small set of repeatable mistakes. The table below is the field guide. Each row is a common pitfall, why it misleads, and how a proper test avoids it.

MOS pitfallWhy it misleadsHow to do it right
Too few ratersA handful of listeners produces a noisy average with a wide confidence interval, so small differences are meaninglessUse enough raters per clip (often 15 or more) and report a confidence interval, not just the mean
Cherry-picked clipsTesting only the sentences a voice reads well hides its failures on numbers, names, and long turnsSample clips at random from your real scripts, including the hard cases the agent must handle
Studio-quality audioRating clean, high-bitrate audio flatters a voice that will actually be heard over a compressed phone lineScore the audio at the codec, bitrate, and packet loss your callers will hear in production
Cross-study comparisonComparing one vendor's number to another's assumes identical listeners and setup, which never holdsRun every voice through one test, with the same raters, clips, and playback, on the same day

The through-line is control. A MOS is trustworthy when one team fixes every variable except the voice under test. The moment you compare across studies, or let the vendor pick the clips, the number stops meaning what you think it means.

How to run a MOS listening test

Here is a concrete, repeatable process for measuring a mean opinion score for voice agents. It works in-house or through an independent evaluator.

1. Define the samples. List the voices you are comparing. Pull a random sample of sentences from your real scripts, including hard cases: dollar amounts, dates, names, and long turns. Aim for enough clips per voice to cover the range, not just easy lines.

2. Render at production quality. Generate each clip through the same audio path your callers hear. Match the codec, bitrate, and any compression of your phone channel. Do not test studio audio you will never ship.

3. Recruit enough raters. Get at least fifteen listeners per clip for a stable average. Use listeners who match your caller base where naturalness perception matters, such as native speakers of the target language.

4. Write clear instructions. Tell raters to score naturalness on the 1 to 5 scale: 5 excellent, 1 bad. Give one worked example. Keep the scale definition on screen during rating.

5. Randomize everything. Shuffle clip order per listener so no voice always plays first. Hide which voice is which. A randomized experiment design stops order and brand bias from leaking into the scores.

6. Add control clips. Insert a few known-good and known-bad anchor clips. They calibrate the scale and let you catch raters who click without listening.

7. Collect one rating per clip per listener. Use absolute category rating: one clip, one score, no side-by-side comparison. Record every individual rating, not just the running average.

8. Compute the MOS and its spread. Average the ratings per voice. Report a confidence interval around each mean so you know whether two voices really differ or just look different.

9. Check rater agreement. Measure whether listeners agreed using inter-rater reliability. Low agreement means the number is shaky, and you should add raters or tighten instructions before you trust it.

Follow these steps and you get a number you can defend. Skip the randomization or the confidence interval, and you get a number that looks precise but hides its own noise.

Testing on your own scripts and phone-quality audio

A MOS from a vendor's demo reel tells you how their voice sounds reading their sentences in their studio. It tells you almost nothing about how it will sound reading your sentences to your callers.

Your scripts are the test that matters. A voice that nails smooth marketing copy can still stumble on a spoken account number or a hyphenated last name. Naturalness is uneven across content. The only way to catch that is to score the voice on the lines it will actually say. Our guide on pronunciation for voice agents covers the specific cases that break otherwise-good voices.

Phone-quality audio matters just as much. Most voice agents run over a phone channel, not a studio link. That channel compresses the audio, drops packets, and narrows the frequency range. A voice rated 4.6 on clean audio can fall well below that once it is squeezed through a real call. Rate the audio your callers hear, not the audio the vendor renders.

This is the difference between a demo score and a deployment score. The demo score is the vendor's best case. The deployment score is what your callers will judge you on. Only one of them predicts how the agent will land in production.

Pairing subjective MOS with objective checks

MOS answers one question well: does this voice sound natural to people? It does not answer every question about voice quality. So a good evaluation pairs the subjective MOS with objective, automatic checks.

Objective checks measure things a number can capture without a listening panel. Does the voice say the right words, or does it drop and mangle them? Are dollar amounts and dates pronounced correctly? Is the timing right, or does it rush and clip? These are checkable against the intended script and do not need a human rating each time.

The two layers cover each other's blind spots. MOS can rate a voice as pleasant while it quietly mispronounces a product name on every call. An objective pronunciation check catches that, even when it does not dent the naturalness score. Run both, and you see both the feel and the facts of the audio. We go deeper on the objective side in our piece on evaluating voice cloning for voice agents.

The split also saves money. Subjective MOS is expensive, because it needs people. So run it periodically, on a sampled set. Run the cheap objective checks continuously, on every build. When an objective check flags a regression, that is your signal to run a fresh listening test. This layering is part of the wider approach in our voice agent evaluation framework.

Why an independent evaluator should run your MOS test

MOS is unusually easy to shade in your own favor. Pick friendly listeners, easy clips, and clean audio, and the number climbs. Because the metric is subjective, none of those choices leave an obvious fingerprint on the final score.

That is exactly why the party reporting a MOS should not be the party that built or sold the voice. The incentive runs the wrong way. A vendor wants the flattering number, and the method that produces it is invisible in the result.

Evalgent is an independent, third-party evaluator built for this. It runs the listening test on your scripts and your phone-quality audio, with a fixed and disclosed method, so the MOS is comparable across every voice you are judging. Because the scoring is external, the number holds up to a buyer, a regulator, or an executive who wants the real picture. We make the broader case in independent voice AI evaluation.

An illustrative example shows the stakes. Suppose a vendor reports a 4.6 MOS from a studio demo. Re-run the test on your own scripts, over your phone codec, with enough raters, and the same voice comes back at 3.8. That 0.8 gap is the difference between a demo you trust and a call your customers actually hear. The figures here are illustrative, not measured, but the direction is why the method matters more than the headline.

Frequently asked questions

What is a good mean opinion score for a voice agent?

A good MOS for a voice agent is usually above 4.0 on the 1 to 5 scale, where 5 is excellent and 4 is good. But a raw number means little without the method. A 4.0 measured on your own scripts over a phone channel beats a 4.6 from a vendor's studio demo, because it reflects real listening conditions.

How many listeners do I need for a MOS test?

Aim for at least fifteen raters per clip to get a stable mean opinion score. Fewer raters produce a noisy average with a wide confidence interval, so small differences between voices become meaningless. More raters tighten the estimate. Always report the confidence interval alongside the mean so you know whether two voices truly differ.

Why can't I compare MOS scores from two different vendors?

You cannot compare MOS across studies because the score has no fixed anchor. Listeners, clips, instructions, and playback devices all shift the number, and none of that is visible in the result itself. The same voice earns different scores in different setups. To compare fairly, run every voice through one test with identical conditions.

What is absolute category rating in a MOS test?

Absolute category rating is the standard MOS method where listeners rate one clip at a time on the 1 to 5 scale, with no reference clip. Each rating judges the clip on its own, not against another. This absolute framing, defined in ITU-T P.800, answers how good a clip sounds rather than which of two clips is better.

Should I measure MOS on studio audio or phone audio?

Measure MOS on the phone-quality audio your callers actually hear. Most voice agents run over a compressed phone channel that drops packets and narrows the frequency range. A voice rated 4.6 on clean studio audio can fall well below that on a real call. Rating studio audio flatters a voice you will never ship in that form.

How is MOS different from objective voice quality metrics?

MOS is subjective: it averages human opinions of naturalness on a 1 to 5 scale. Objective metrics are automatic and measure checkable facts, like whether words and numbers are pronounced correctly. MOS captures how a voice feels; objective checks capture what it gets right. Pair them, because a pleasant voice can still mispronounce a product name every call.

How do I know if my MOS number is reliable?

A reliable MOS comes with a confidence interval and high inter-rater reliability. The confidence interval shows how much the mean might move with different raters. Inter-rater reliability shows whether listeners agreed. If agreement is low or the interval is wide, the number is shaky. Add raters or tighten instructions before you act on it.

Who should run a voice agent's MOS listening test?

An independent, third-party evaluator should run the MOS test, because the builder or vendor has a reason to report the flattering number. Subjective scoring is easy to shade with friendly listeners and easy clips. An external party like Evalgent uses a fixed, disclosed method on your own scripts, so the result is comparable and defensible to buyers and regulators.

The bottom line

A mean opinion score is only trustworthy when you run the listening test yourself, on your own scripts and phone-quality audio. A vendor's number from a different study sits on a different ruler and cannot be compared to yours.

Want a mean opinion score you can defend, measured on your calls with a disclosed method? Book a demo and see how Evalgent runs MOS as an independent third party.

Related Articles