Menu modifiers in voice ordering: how to test food ordering voice agents

On this page
The hard part of voice ordering is not "a cheeseburger." It is "a cheeseburger, no onions, extra pickles, make it a large combo with a Diet Coke, light ice, and actually make the fries a side salad with ranch." Every clause in that sentence is a modifier, and each one has to land on the right item, in the right group, with the right quantity, in a payload the point-of-sale (POS) system will accept.
The independent data points the same way. In Intouch Insight's January 2025 mystery-shop study of voice AI drive-thrus, AI lanes got the order right 83% of the time against 89% for traditional lanes, and 65% of the errors were tied to customizations. The 2026 QSR Drive-Thru Report found AI order accuracy improved from 84% to 89% year over year, but non-AI accuracy at the same restaurants was 95.7%, and customization was again called the "Achilles heel."
This guide is written for the engineer who builds the ordering agent in-house, usually on LiveKit or Pipecat with a hosted STT, an LLM with tool calls, and a POS integration. It covers what modifiers are, how POS menus represent them, the speech failures that break them, a test matrix and scoring formulas, a readiness checklist, vendor questions, and an order-diff scorer you can adapt.
What are menu modifiers in voice ordering?
Menu modifiers are the configurable options attached to a menu item that change what the kitchen makes or what the guest pays: a size choice, a topping added or removed, a sauce on the side, a dressing for a side salad, or a drink upsized in a combo. In a POS, modifiers live in modifier groups that carry rules: whether a choice is required, how many options can be picked, and which options nest under others.
For a voice agent, every modifier is a slot the LLM must fill from speech, attach to the correct line item, and validate against those group rules before the order is submitted. That is why modifiers, not items, dominate the error budget.
Types of menu modifiers voice AI must handle
| Modifier type | Example utterance | Group rule | What the agent must do | Typical failure |
|---|---|---|---|---|
| Required single-select | "A large" | min 1, max 1 | Ask if missing; reject a second value | Never asks; defaults silently |
| Optional multi-select | "Add bacon and jalapeños" | min 0, max N or none | Accept several; enforce max | Drops the second option |
| Negation ("no X") | "No onions" | Pre-modifier on a default | Remove a default ingredient | Heard as "an onion"; ignored |
| Intensity ("extra / light X") | "Extra cheese, light ice" | Pre-modifier, often priced | Apply and price correctly | Mapped to plain "cheese" |
| Nested | "Side salad with ranch" | Child group required when parent picked | Ask the child question | Salad sent without dressing |
| Quantity on modifier | "Double cheese" | Same option twice or quantity field | Encode the POS-specific way | Single cheese |
| Quantity split | "Two burgers, one without pickles" | Separate lines when modifiers differ | Split into two lines | Both without pickles |
| Substitution | "Sub the fries for onion rings" | Swap within a group, may cost more | Replace, not add | Both sides added |
| Half-and-half | "Half pepperoni, half mushroom" | Portions | Assign toppings to halves | Halves swapped or whole pizza |
| Combo build and upsize | "Make it a large meal" | Item becomes a bundle with its own groups | Convert item, fill side and drink | Upsize applied to the burger only |
| Free-text request | "Sauce on the side, cut in half" | Special request text | Pass through verbatim | Lost, or invented as a modifier |
The PIZZA paper from Amazon's Alexa AI team, a benchmark for parsing pizza and drink orders, makes the same point formally: food orders have semantics that "cannot be captured by flat slots and intents." Its example "three pizzas, one with peppers and the others with ham" needs the parser to split one quantity into two differently modified groups, which is exactly the quantity-split row above (Arkoudas et al., 2022).
How POS menu APIs model modifiers
Your agent does not submit English. It submits a structured order, and the structure is set by the POS. Understanding the schema tells you where the edge cases hide.

Toast. In Toast's menus API, a ModifierGroup carries `minSelections` and `maxSelections`. `minSelections` of 0 means optional, 1 or more means required, and `maxSelections` is null when the guest can pick an unlimited number. Pre-modifiers such as NO and EXTRA are separate objects referenced from the group. Toast's orders API guide on modifiers documents four behaviors that matter for voice:
1. Modifier quantity must equal item quantity. Two cheeseburgers with bacon means the bacon modifier has quantity 2. One with bacon and one without must be two separate selections. Mismatched quantities are rejected.
2. Default modifiers are not added automatically via the API. If the turkey sandwich defaults to lettuce and tomato, your payload has to include them. Omitting a default tells the kitchen to leave it off. An agent that forgets defaults strips every sandwich, and nobody said "no."
3. Doubling is repetition. "Double cheese" is the cheese modifier added twice.
4. Halves are portions, and free text is a special request. Half-and-half pizzas use `selectionType` `PORTION`; "dressing on the side" can go in as `SPECIAL_REQUEST` with up to 1,000 characters, and it only shows on the kitchen display if it sits inside `modifiers`.
Toast's pre-modifier settings also let a restaurant set a price multiplier, for example 0 for NO and 2 for EXTRA. So "extra bacon" is a pricing event, and a wrong pre-modifier is a wrong total.
Square. In the Square Catalog API, a CatalogModifierList uses `min_selected_modifiers` and `max_selected_modifiers`; the older `selection_type` field (SINGLE or MULTIPLE) is deprecated. A default of -1 means no minimum or no maximum, and both can be overridden per item. There is also a beta `modifier_type` of `TEXT` for free-text modifiers and a beta `allow_quantities` flag for picking the same modifier more than once.
Online ordering platforms such as Olo use comparable option-group concepts. Don't assume field names carry across systems. Read the schema for the system you integrate with, and keep a canonical internal order format so tests don't depend on one vendor's JSON.
The testing lesson: the same spoken order has several valid encodings (two lines of 1 or one line of 2; defaults listed or implied), and a few invalid ones the POS rejects or, worse, accepts with the wrong meaning. Your scorer has to canonicalize before it compares, or it will report false errors and miss real ones.
Speech-specific failures in food ordering
Speech adds failure modes before the LLM ever sees the order. Map each one to a stage so you know which assertion catches it.

Homophones and number words
"Two" and "to," "four" and "for," "a" and "eight," "no" and "an": quantities and negations are short, unstressed function words, which makes them easy to mishear and expensive to get wrong. "Can I get to large fries" is ambiguous in text; "I want no onion" heard as "I want an onion" flips the meaning. Measure quantity accuracy and negation accuracy as their own metrics. Word error rate averages them away, which is why entity-level STT accuracy matters more than WER here.
Menu names are the other half. Deepgram's own keyterm prompting docs use drive-thru examples ("nacho" heard as "macho," "bacon" as "bake in") and support up to 100 keyterms within a 500-token limit, recommending 20 to 50. Two gotchas from that page: separating terms with commas does not error, it silently boosts nothing, and on Nova-3 and Flux streams a `Configure` message replaces the whole keyterm list, so you can swap menu sections mid-call but must resend everything you still need. A 300-item menu will not fit. Pick the confusable items, the limited-time items, and the brand names, and test recall on exactly those.
Corrections and cart edits
Callers revise: "actually make that a medium," "scratch the fries," "the second one without pickles." Each correction is a reference to an earlier line, and the agent has to resolve which line. The SpokenWOZ benchmark (NeurIPS 2023, 249 hours of human-to-human spoken task dialogues including restaurant booking) found the best dialogue state tracker reached only 25.65% joint goal accuracy, with cross-turn slots called out as a new challenge, which is what an order revised over several turns looks like. Your test set needs corrections at the start, middle, and end of the order, and assertions on the cart after every turn, not only at checkout. The failure pattern is covered in our guide to testing memory and context.
Long orders have their own failure. In the PIZZA paper's error analysis, the best model (BART generating executable representations, 78.56% exact match on the test set) missed slots and dropped whole sub-orders on lengthy multi-item requests that a hand-written grammar parsed correctly. Your test set should include five-item orders spoken in one breath, not only one-item turns.
Endpointing, silence, and the menu board
At a drive-thru, the guest often stops talking to read the board. A short endpointing timer ends the turn mid-order ("I'll have the..."), and the agent answers a half-sentence. A long one adds dead air at every turn. Test pauses of several lengths inside an item name and between items; our comparison of end-of-turn detection options covers the trade-off.
Silence is also a hallucination risk for some STT models. The Careless Whisper study (FAccT 2024) found about 1% of Whisper transcriptions contained fabricated phrases with no basis in the audio, and the hallucinations varied from run to run. If an idle-engine pause can produce text, that text can become a line item. Include long non-speech segments in your noise set and assert that nothing gets added to the cart.
Noise, second speakers, and read-back
Drive-thru audio carries engines, wind, and passengers; phone orders carry kitchen clatter and TV in the background. A passenger saying "and a shake" may or may not be part of the order, and the agent must not add it silently. Test a second voice at lower level and at equal level, and assert that the agent either ignores it or confirms it. Run a structured background-noise protocol with SNR tiers rather than one "noisy" clip.
Read-back is the safety net, but it is weaker than it looks. The 2026 QSR report found that among 70 AI orders where the system confirmed the order back to the guest, 7 were still incorrect. Separately, across all orders at locations without a usable confirmation board, accuracy was 92% when the order was repeated back versus 76% when it was not. Confirmation helps, but only if what is read back is the payload that will be submitted, every modifier included, and the guest has a real chance to correct it.
Drive-thru vs phone ordering: what changes in testing
| Dimension | Drive-thru | Phone ordering |
|---|---|---|
| Audio path | Outdoor speaker post and mic; quality depends on hardware | PSTN or SIP, usually narrowband 8 kHz; codec artifacts |
| Noise profile | Engines, wind, traffic, adjacent cars | Kitchen, TV, street, car speakerphone |
| Speakers | Driver plus passengers, children in the back | Usually one caller; speakerphone adds others |
| Menu visibility | Guest reads names off the board, often exact names | Guest recalls items; more vague or off-menu names |
| Order size | Mostly a few items per car | Can be large (family or office orders), plus address and payment |
| Confirmation | Visual confirmation board plus voice | Voice-only read-back; nothing to look at |
| Handoff | Crew member on headset can take over | Transfer to a person on a busy line, or callback |
| Latency tolerance | A car is waiting; pauses are visible | Silence on a phone line reads as a dropped call |
Two consequences for testing. First, phone orders need a phone audio quality pass with real codecs, because a test set recorded at 16 kHz on a laptop overstates accuracy. Second, drive-thru accuracy has to be reported separately for orders the AI completed and orders that were handed to a person. The 2026 QSR report found AI handled the full interaction in 40% of AI engagements, up from 20% the year before. An accuracy number that silently excludes handoffs is not comparable to one that includes them.
Order accuracy metrics and formulas
Score the POS payload, not the transcript and not the agent's description of the order. Use one canonical format: every line expanded to quantity-1 units, each unit described by its item ID plus the sorted set of (parent path, portion, group, option, pre-modifier) tuples. Then define:
- Order exact-match accuracy (OEA) = orders whose canonical payload equals the canonical gold order ÷ total orders. This is the number the guest experiences. One wrong modifier fails the order.
- Item precision = matched item units ÷ item units in the payload. Item recall = matched item units ÷ item units in the gold order. Low precision means phantom items; low recall means dropped items.
- Modifier exact-match rate (MEM) = item units whose full modifier set matches gold ÷ item units whose item ID matched. This isolates modifier errors from item errors.
- Misattachment count = modifiers missing from one item and present on another in the same order. "No onions" moved from the burger to the fries is the classic case, and it is the error a transcript review misses.
- Required-group completion rate = required groups that were filled by the guest's words or by an explicit question ÷ required groups in the order. Silent defaults count as failures.
- Correction success rate = spoken corrections reflected correctly in the next cart state ÷ spoken corrections.
- Read-back confirmation rate = submitted orders preceded by a full read-back ÷ submitted orders. Pair it with read-back fidelity = read-backs whose content equals the submitted payload ÷ read-backs, and post-read-back error rate = wrong orders that were read back ÷ orders read back. The QSR data point above would be a 10% post-read-back error rate.
- Handoff rate = orders where a person took over ÷ orders started. Define "person" broadly; see the vendor section below.
Severity matters. A missed "no" on an ingredient the guest asked to remove is worse than a missed "light ice." Tag negation errors as critical and report them separately, even when OEA looks fine.
How many test orders you need
Accuracy is a proportion, so its margin of error shrinks slowly. With 200 orders at 90% OEA, the 95% confidence interval is about ±4.2 points (1.96 × √(0.9 × 0.1 / 200) ≈ 0.042). That means a 3-point change between two builds is noise at that sample size.
To detect a drop from 95% to 90% with 80% power at α = 0.05, a two-proportion test needs about (1.96 + 0.84)² × (0.95 × 0.05 + 0.90 × 0.10) ÷ 0.05² ≈ 431 orders per arm. For release gates, run a fixed critical set many times and watch for any new failure on it, rather than relying only on a single aggregate number. Our guide to building a golden dataset covers how to keep that set stable.
A test matrix for food ordering voice agents
Build the suite as modifier classes crossed with conditions, not as a list of happy-path orders.

The worked example in the diagram uses 10 modifier classes × 5 conditions = 50 cells. Write 3 phrasings per cell (menu-board wording, casual wording, and a regional or accented variant) and run each 3 times, because LLM-driven agents are not repeatable run to run. That is 450 simulated orders for a full pre-launch pass, and 180 for the 20 critical cells you run on every release.
Fill in the cells with your own menu: the items with the most modifier groups, the items with defaults guests often remove, your limited-time offers, and any item whose name is close to another. Add three cross-cutting scenario sets on top:
- Menu drift: an item 86'd since the last menu sync, a new item added this morning, a price change. Assert graceful handling and no invented items.
- POS failure: a 4xx from the order endpoint, a timeout, a quantity-rule rejection. Assert the agent does not tell the guest the order is placed. See detecting silent tool failures.
- Upsell: an offer accepted, declined, and half-heard. Assert accepted upsells land in the payload and declined ones don't.
For tool-call structure (argument types, enum values, nested arrays), our tool calling test cases include food-ordering examples you can copy, and tool argument accuracy explains how to score them.
An order-diff scorer you can adapt
The scorer below is simplified and illustrative, but it runs as written (standard-library Python 3). It canonicalizes both orders into quantity-1 units, so "two burgers, one without onions" compares equal whether it was encoded as one line or two, then reports exact match, item precision and recall, modifier exact-match rate, and misattached modifiers.
from collections import Counter
from dataclasses import dataclass, field
@dataclass(frozen=True)
class Mod:
group: str # modifier group id, e.g. "toppings"
option: str # option id, e.g. "onion"
pre: str = "" # pre-modifier: "", "NO", "EXTRA", "SIDE", "LIGHT"
portion: str = "whole" # "whole", "half1", "half2"
children: tuple = () # nested Mods, e.g. dressing under a side salad
@dataclass
class Line:
item: str
qty: int
mods: tuple = field(default_factory=tuple)
def flat(mods, path=""):
"""Flatten a modifier tree into hashable tuples that keep their parent path."""
out = []
for m in mods:
out.append((path, m.portion, m.group, m.option, m.pre))
out += flat(m.children, path + "/" + m.option)
return out
def canon(lines):
"""Expand every line into qty-1 units so equivalent encodings compare equal."""
units = []
for ln in lines:
sig = (ln.item, tuple(sorted(flat(ln.mods))))
units += [sig] * ln.qty
return Counter(units)
def score(gold, pred):
g, p = canon(gold), canon(pred)
gi, pi = Counter(), Counter()
for (item, _), n in g.items(): gi[item] += n
for (item, _), n in p.items(): pi[item] += n
item_tp = sum((gi & pi).values())
unit_tp = sum((g & p).values()) # units identical including all modifiers
gm = Counter((i, m[2:]) for (i, ms), n in g.items() for m in ms for _ in range(n))
pm = Counter((i, m[2:]) for (i, ms), n in p.items() for m in ms for _ in range(n))
lost = {m for (_, m) in (gm - pm)}
gained = {m for (_, m) in (pm - gm)}
return {
"order_exact": g == p,
"item_precision": item_tp / max(sum(pi.values()), 1),
"item_recall": item_tp / max(sum(gi.values()), 1),
"modifier_exact_rate": unit_tp / max(item_tp, 1),
"misattached_mods": sorted(lost & gained),
"missing_or_wrong": list((g - p).items()),
"extra_or_wrong": list((p - g).items()),
}
gold = [
Line("cheeseburger", 1, (Mod("toppings", "onion", "NO"),)),
Line("cheeseburger", 1),
Line("fries", 1, (Mod("size", "large"),)),
]
pred = [
Line("cheeseburger", 2), # "no onion" lost...
Line("fries", 1, (Mod("size", "large"), Mod("toppings", "onion", "NO"))), # ...and moved here
]
print(score(gold, pred))
# order_exact False, item_precision 1.0, item_recall 1.0,
# modifier_exact_rate 0.333..., misattached_mods [('toppings', 'onion', 'NO')]Note what this example shows: item precision and recall are perfect, so an item-level dashboard would call the order correct. Only the modifier-level metrics catch it.
Three extensions to make before you rely on it. First, map POS-specific payloads (Toast selections, Square line items) into `Line` and `Mod` with an adapter, and resolve default modifiers explicitly, so an omitted default shows up as a removal. Second, add a severity table keyed on (group, pre-modifier) so a lost NO is critical. Third, run the same scorer on the cart after each turn to compute correction success rate, not only on the final payload.
How to test a food ordering voice agent before rollout
1. Export the live menu and rules. Pull items, modifier groups, min/max selections, defaults, pre-modifiers, and portions from the POS API, and snapshot it with a version so tests can pin a menu.
2. Write gold orders as structured data. For each test call, write the expected canonical order first, then the caller script that should produce it. The gold order is the contract.
3. Fill the matrix. Cover the 10 modifier classes across clean, phone, drive-thru noise, second speaker, and correction conditions, prioritizing the critical cells.
4. Generate realistic audio. Use scripted callers with varied pace, accents, and disfluencies, mix recorded noise at defined SNR tiers, and pass phone tests through a real narrowband codec path. Our accent robustness guide covers caller diversity.
5. Capture every layer. Log the transcript, each tool call, the cart after each turn, the read-back text, and the final POS request and response for every test call.
6. Score at the payload, diagnose upstream. Run the order-diff scorer on the submitted payload, then use transcript and cart logs to attribute each failure to audio, STT, cart logic, read-back, or POS.
7. Repeat runs and read variance. Run each scenario at least 3 times; a scenario that passes 2 of 3 is a failing scenario.
8. Set gates and re-run on every change. Gate releases on OEA, negation errors, misattachments, and read-back fidelity, and re-run the critical set on every prompt, model, STT, or menu change.
9. Pilot with shadow scoring. In a limited rollout, score a sample of real orders against what the kitchen actually made or what the guest corrected, and compare against your offline numbers.
Voice ordering readiness assessment checklist
Use this before any live traffic. Every "no" is a known risk you are choosing to ship.
| Area | Readiness question | Ready when |
|---|---|---|
| Menu data | Does the agent read modifier rules from the POS rather than a hand-copied prompt? | Menu is synced and versioned |
| Required groups | Does the agent ask for every required choice the guest omits? | Required-group completion ≥ 99% on the test set |
| Defaults | Are default modifiers included in payloads and removed only when asked? | Zero unintended removals |
| Negation | Are "no X" requests captured and placed on the right item? | Zero critical negation errors in the critical set |
| Quantity splits | Do "two, one without" orders produce separate lines? | All quantity-split cells pass 3 of 3 runs |
| Corrections | Do mid-order changes update the right line? | Correction success ≥ 95% |
| Noise | Is accuracy measured per SNR tier and per channel? | Drive-thru and phone numbers reported separately |
| Read-back | Is the read-back generated from the payload to be submitted? | Read-back fidelity 100% |
| POS errors | Does the agent detect rejected or timed-out submissions? | No "order placed" on failure |
| Handoff | Is there a clean handoff path, and is it measured? | Handoff rate tracked with reasons |
| Regression | Does every change re-run the critical set? | Gate in CI |
These thresholds are suggested starting points, not industry standards. Set your own from what a wrong order costs you.
Voice ordering vendor evaluation questions
If you are buying instead of building, or comparing an in-house build to a vendor, the hard part is that vendors define their metrics differently. The SEC's 2025 order against Presto Automation is a useful case study: the company reported "automated order completion" and "non-intervention" rates of 95% or more, but those rates measured orders completed without restaurant staff involvement, not without any human involvement, and off-site human agents processed most orders (SEC order, January 2025). The order also says the more advanced pilot version needed a human agent to enter orders about 70% of the time from June to December 2023.
Ask every vendor:
1. What counts as an intervention? Restaurant staff only, or any human, including remote agents who review or correct orders?
2. What is the accuracy denominator? All orders started, or only orders the AI completed without handoff?
3. What is "accurate"? Exact match on every item and modifier, or "mostly right"? Who judged it, and against what?
4. Is accuracy measured on the POS payload or on the transcript?
5. What are modifier-level and negation error rates, separately from item accuracy?
6. How does accuracy vary by channel and noise? Drive-thru, phone, and peak hours, reported separately.
7. How do you handle default modifiers, quantity splits, halves, and nested groups in our POS? Ask for payload examples from a test order.
8. What happens when the POS rejects an order or an item is 86'd mid-call?
9. How fast do menu changes reach the agent, and how are they tested?
10. Can we run our own gold orders through your system and score the payloads ourselves?
The last question matters most. A vendor that can only show you its own dashboard is asking you to trust its definitions. Our vendor evaluation guide, the order status vendor evaluation, and the POC bake-off guide cover how to run that comparison with the same test cases for every candidate. This is also where an independent evaluator like Evalgent fits: the same gold orders and the same scorer, applied to every vendor and to your own build.
Frequently asked questions
What are menu modifiers in voice ordering?
Menu modifiers are the options that change a menu item: size, added or removed toppings, sauces, nested sides, halves, and combo upgrades. In a POS they live in modifier groups with rules for required choices and how many options a guest can pick. A voice agent must fill each one from speech and attach it to the right item.
What types of menu modifiers must voice AI handle?
At minimum: required single-select (size), optional multi-select (toppings), negation ("no onions"), intensity ("extra cheese," "light ice"), nested groups (side salad then dressing), quantity on a modifier (double cheese), quantity splits ("two, one without"), substitutions, half-and-half portions, combo builds and upsizes, and free-text special requests.
How do I test AI voice agents for drive-thru ordering before rolling out?
Write structured gold orders, generate caller audio with drive-thru noise, second speakers, pauses, and corrections, and run each scenario several times. Score the submitted POS payload against the gold order at item and modifier level. Gate on exact-match accuracy, negation errors, misattachments, and read-back fidelity, then pilot with shadow scoring.
How is order accuracy measured for voice ordering?
Use order exact-match accuracy: the share of orders whose canonical POS payload equals the gold order, item and modifier included. Add item precision and recall, modifier exact-match rate, misattachment count, required-group completion, correction success, and read-back fidelity. Score payloads, not transcripts, because a correct transcript can still produce a wrong ticket.
Why do voice ordering agents get modifiers wrong?
Modifiers are short, unstressed words that STT mishears ("no" vs "an," "two" vs "to"). They must attach to the right item across several turns, follow group rules, and be encoded the POS-specific way, such as including default modifiers or splitting lines when quantities differ. Each step is a separate failure point.
What is a voice ordering readiness assessment?
It is a pre-launch check that the agent reads menu rules from the POS, asks for required choices, handles defaults, negations, quantity splits and corrections, holds accuracy under realistic noise, reads back the exact payload, detects POS failures, and is re-tested on every change. Each item needs a measurable pass condition.
How should you evaluate voice ordering vendors?
Ask how they define intervention, what the accuracy denominator is, whether accuracy is scored on the POS payload, and what their modifier-level and negation error rates are by channel. Then run your own gold orders through each vendor and score the payloads yourself with one scorer, so every candidate is measured the same way.
Does reading the order back guarantee accuracy?
No. Read-back helps, but in the 2026 QSR Drive-Thru Report, 7 of 70 AI orders that were confirmed back to guests were still wrong. Read-back only works if it is generated from the exact payload to be submitted, includes every modifier, and gives the guest a real chance to correct it.
The bottom line
Food ordering voice agents rarely fail on the item; they fail on the modifier, the quantity split, the dropped default, and the correction applied to the wrong line. Test them by scoring the POS payload against structured gold orders across modifier classes, noise, second speakers, and corrections, and treat any vendor accuracy number you cannot reproduce on your own orders as unverified.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more