Why Voice Agents Get Stuck Repeating Themselves (and How to Detect Repetition Loops)

On this page
The agent asks for a date of birth. The caller gives it. The agent says "Sorry, I didn't get that. Can I get your date of birth?" The caller says it again, slower and louder. Same answer. By the fourth ask, the caller is shouting "representative" or hanging up.
Search results for this are GitHub issues and forum threads, each describing one cause. In practice, at least eight mechanisms produce the same symptom, in different layers of the stack. A frequency penalty won't fix a date validator. A better STT model won't fix a Pipecat node that never transitions.
This guide sorts the causes by layer, with a transcript signature for each. It pulls in research on dialogue error recovery, correction prosody and neural text degeneration. Then it gives you four detection signals, a re-ask rate formula, a Python detector for transcript JSON, a fix ladder and a test matrix that reproduces each cause.
What a repetition loop is
A repetition loop happens when the agent produces the same prompt, or the same meaning, on two or more agent turns while the conversation state stays put. The wording can change. "Can I get your date of birth?" and "Could you tell me when you were born?" are the same loop. The test is whether anything moved: a slot filled, a node changed, a tool returned something new.
There are three families, and they need different fixes:
- Re-ask loops. The agent asks for the same piece of information again because it doesn't have it, or thinks it doesn't. This is the most common family on scheduling, intake, verification and payments calls.
- Degenerate repetition. The LLM repeats a sentence or a whole response even though the state moved. This is a decoding problem, not a data-capture problem.
- Echo duplicates. The agent answers the same thing twice because it got two user turns for one intent. Usually the caller repeated themselves during a long pause.
A handoff loop is a cousin. There the cycle runs across agents or departments. Here it runs inside one agent, often inside one slot.
The eight causes, by layer
Each cause leaves a different trace in your logs. The table is the triage map. The sections below explain the mechanism behind each row.
| # | Cause | Layer | Transcript signature | Where to look |
|---|---|---|---|---|
| 1 | STT never captured the answer | Audio / STT | User turn empty, cut off or wrong; agent re-asks | Endpointing events, final vs interim transcripts |
| 2 | Validator rejects valid input | Tool / code | User turn contains a correct answer; tool returns invalid | Tool args vs raw transcript |
| 3 | Tool failed, retried silently | Tool | Tool error or timeout, then the LLM asks again | Tool outputs, latency spikes |
| 4 | Flow or task not advancing | Orchestration | Slot filled in tool output, node unchanged | Node/task transition log |
| 5 | Context dropped the answer | Context | Answer exists earlier in the call, missing from LLM input | Chat context size, handoffs, resets |
| 6 | Interruption truncated the agent | Turn-taking | Agent turn marked interrupted; question re-asked | Interrupted flags, overlap timestamps |
| 7 | LLM degenerate repetition | LLM decoding | Near-identical agent turns even after progress | Sampling settings, repeated text in context |
| 8 | Caller repeats due to latency | Latency | Two user turns with the same intent; two agent replies | Response latency before the repeat |

1. STT never captures the answer
The agent can only re-ask if it thinks it has nothing. The most common reason is that the transcript really is empty or broken. Three paths lead there.
First, endpointing cuts the turn early. The caller says "March third," pauses to think, then says "nineteen eighty-two." A short silence timer commits "March third" as the final turn. The validator rejects a date without a year, and the agent re-asks. The caller's second answer often has the same pause, so it fails the same way. If your stack uses a VAD silence timer rather than a semantic end-of-turn model, test pauses inside numbers first. The end-of-turn detection comparison covers how each method treats mid-utterance pauses.
Second, noise and narrowband audio mangle entities. One wrong digit turns a valid account number into an invalid one. See STT entity accuracy and background noise testing for how to measure this separately from WER.
Third, the STT invents text. Koenecke et al. (2024) found that roughly 1% of Whisper transcriptions in their study contained entire hallucinated phrases that weren't in the audio. Hallucinations were more common for speakers with longer non-vocal stretches. On a call, a hallucinated phrase during silence becomes a user turn the agent must respond to, often with "Sorry, could you repeat that?"
2. Slot validation rejects valid input
This one is pure code, and it hides well because the transcript looks fine. The caller says "March third, eighty-two." The STT returns `March 3rd 82` or `3/3/82` depending on formatting options. The LLM passes the string to a tool. The tool expects `YYYY-MM-DD` and returns `invalid_date`. The LLM does what any polite assistant would do and asks again.
Spoken forms that commonly break naive validators:
| Slot | Spoken form | What a strict parser often expects |
|---|---|---|
| Date of birth | "March third eighty-two", "three three eighty-two", "the third of March" | `1982-03-03` |
| Phone | "five five five, twelve twelve", "oh" for zero | 10 digits |
| "john dot smith at gmail dot com" | `john.smith@gmail.com` | |
| Account / member ID | "double four seven", "A as in apple" | `447`, `A` |
| Money | "a hundred and twenty-five fifty" | `125.50` |
The fix is to normalize spoken forms before validating, and to put the raw transcript in the tool args next to the parsed value. LiveKit's prebuilt `GetDOBTask`, `GetPhoneNumberTask` and `GetEmailTask` exist because this is hard. The docs describe them as handling spoken dates, normalizing spoken digits and coping with noisy transcription. Even if you don't use them, test your own parser against the spoken forms above.
3. A tool failure retried silently
The tool times out or returns a 500. The framework hands the error to the LLM as the tool result. The LLM doesn't say "our system is down." It apologizes and asks for the account number again, because from its point of view the input might have been wrong. The caller repeats a correct number into a broken system.
The signature is a tool error or timeout immediately before a re-ask, with an unchanged argument on the retry. If the same argument fails twice, the caller's input isn't the problem. Detecting silent tool failures covers how to surface these as errors instead of letting the LLM paper over them.
4. The flow or task never advances
In a structured flow, the next question is decided by code, not the LLM. If the transition doesn't fire, the agent stays in the node whose instructions say "ask for the date of birth," so it asks again.
In Pipecat Flows, a handler returns a `(result, next_node)` tuple. The Flows 1.0 migration guide shows that returning `None` as the second element means no transition. That's correct for a lookup that should stay in the node. It's a loop when a branch you forgot to handle, such as a partial match, falls through to `None`. Flows now ships in core Pipecat as `pipecat.flows` (since `pipecat-ai` 1.5.0), so rule out a stale `pipecat-ai-flows` install too.
In LiveKit, an `AgentTask` holds the session until something calls `complete()`. The tasks docs warn that "an LLM might not call a task's completion tool on the first turn." If the completion tool's description is vague, the LLM may keep chatting instead of calling it. `TaskGroup` adds a second path: it lets the LLM regress to earlier tasks "as often as needed" for corrections. A caller who says "wait, no" at the wrong moment can send the group back to a step that was already done.
5. Context drops the answer
The answer was captured, but the LLM can no longer see it. Three documented behaviors cause this in LiveKit. Per the chat context docs, a new agent or task starts with an empty context by default. `truncate(max_items=6)` keeps only the last six items. And `TaskGroup` summarizes its interactions into one message by default when it finishes. In Pipecat Flows, the `RESET` context strategy discards previous history on a node transition.
Together they produce "Can I get your name?" eight minutes after the caller gave it. Keep collected slot values in structured session state, not only in chat history, and re-inject them on every handoff. The memory and context testing guide has the scenarios for this.
6. An interruption truncated the agent
When a caller barges in, LiveKit truncates the agent's history entry to the part the caller actually heard. The turns overview describes this, and the message arrives with `interrupted=True`. That's the right behavior, and it creates two loop paths.
If the cut came before the question, the LLM sees that it never asked, so it asks again. If the caller answered the half-heard question, the answer may not match what the LLM thinks it asked. False interruptions make this much worse, because a cough or a "mm-hm" cuts the agent off mid-question. The LiveKit false interruptions guide covers the settings that drive this, including a `min_words` value that drops short answers like "No."
7. LLM degenerate repetition
This is the only cause that lives in the model itself. Holtzman et al. (ICLR 2020) showed that using likelihood as a decoding objective produces text that is "bland and strangely repetitive," even from strong models. Greedy and beam search are the worst offenders. Their fix, nucleus (top-p) sampling, cut the low-probability tail while keeping enough diversity to avoid loops.
Xu et al. (NeurIPS 2022) explained why loops persist once they start. Consecutive sentence-level repetitions are rare in human text, about 0.02% in Wikitext-103. But models prefer to repeat the previous sentence. The repetition has a self-reinforcement effect: the more times a sentence appears in the context, the more likely the model is to generate it again. Sentences with high initial probability reinforce the most.
That finding matters for voice agents. "Sorry, I didn't catch that. Could you repeat your date of birth?" is a high-probability sentence. Every time it lands in the chat history, the next turn is more likely to produce it again. The loop feeds itself through the context, not only through the decoder.
Two consequences follow:
- Low temperature makes it worse. Teams set temperature near 0 for consistency, which pushes decoding toward the greedy behavior Holtzman et al. measured.
- Frequency penalties don't reach across turns. OpenAI's documented penalty formula adjusts each token's logit by `c[j]`, "how often that token was sampled prior to the current position." That describes repetition inside the response being generated. A sentence from two turns ago sits in the prompt, and the formula as documented doesn't target it. Many realtime and speech-to-speech APIs expose no penalty at all.
The fix that does reach across turns is to edit the context. When a re-ask happens, collapse the earlier identical prompts or add an instruction that names the change: "You've asked for this twice. Ask for month, day and year one at a time."
8. The caller repeats because of latency
Callers fill silence. If the agent takes two or three seconds to answer, many callers say the request again. The STT now produces two user turns. Depending on how the framework handles a new user turn during generation, you get one cancelled reply and one real one, or two full replies that say the same thing. The caller hears the agent repeat itself, but the root cause is the delay before the first reply.
The signature is a user turn whose text closely matches the previous user turn, arriving after a long response gap. The cost of latency post covers how gaps compound on a call.
Why the second ask fails more often than the first
Most agents handle a failed capture with the same move: apologize and ask again. Two classic studies show why that move has poor odds.
Bohus and Rudnicky (SIGDIAL 2005) ran a conference-room reservation system that picked among ten recovery strategies at random after each non-understanding. That gave each strategy a fair comparison. The recovery rate is the share of attempts where the next user turn was understood. The spread was wide:
| Strategy | What the system says | Recovery rate |
|---|---|---|
| MoveOn | Skips the failed question and asks a different one | 64.4% |
| FullHelp | Explains the current state and what to say | 58.5% |
| TerseYouCanSay | Gives a short example of a valid answer | 56.5% |
| Reprompt | Repeats the previous prompt | 49.2% |
| AskRephrase | Asks the user to say it differently | 48.6% |
| AskRepeat | "Can you please repeat that?" | 33.7% |
| Yield | Says nothing, waits | 31.2% |

The paper also confirmed prior findings that an error makes another error on the next turn more likely. The authors describe a "spiral of errors": patience runs out and acoustic mismatches grow. For non-native speakers, 26.3% of turns were non-understandings and the recovery rate was only 39.3%.
Litman, Hirschberg and Swerts (2006) explain part of the spiral. Corrections, the turns where a user tries to fix a system error, tend to be hyperarticulated: louder, slower, higher in pitch. In their corpus, 52% of corrections showed perceptual hyperarticulation, against 12% of other turns. Hyperarticulated corrections were misrecognized 70% of the time, against 52% for other corrections. The prompt shapes the response too. After rejections, when the system said a close paraphrase of "Can you please repeat your utterance?", 59% of user corrections were plain repetitions.
These systems predate neural ASR, so don't carry over the exact rates. The direction holds: a caller asked to "repeat that" repeats it louder and slower, which is speech your STT saw less of in training.
Worked example, with illustrative numbers. Say first-attempt capture works 85% of the time, and each retry works 50% of the time because of the spiral. With three asks before escalation, the share of callers who fail all three is 0.15 × 0.5 × 0.5 = 3.75%. If you wrongly assume retries are independent at 85%, you'd predict 0.15³ = 0.34%. That's an elevenfold underestimate of how many callers end up stuck.
Four signals that detect a repetition loop
No single signal catches every family. Use all four.
1. Similarity between consecutive agent turns. Compute word-trigram Jaccard and a character-level ratio between each agent turn and the previous one, and take the max. This catches verbatim and near-verbatim repeats, family 7 especially. It misses paraphrased re-asks. In the detector below, "Sorry, I didn't get that. Can I get your date of birth?" against "I'm sorry. Can I get your date of birth please?" scores 0.65, under a 0.80 flag. For paraphrase, add embedding cosine similarity, but calibrate it on your own transcripts first.
2. Same-question counter per slot. Tag each agent turn with the slot it asks for. If you use structured flows, the node or task already knows. Otherwise, a small classifier over the agent text works. Count asks per slot per call. This is the most reliable signal for re-ask loops, and it doesn't care about wording.
3. Re-ask rate. At fleet level:
re-ask rate = Σ (asks_for_slot − 1) / number of slots requested
loop-call rate = calls with any slot asked ≥ 3 times / total callsA call that asks for 4 slots and re-asks one of them twice has a re-ask rate of 2 / 4 = 0.5. Track it per slot. A spike on `dob` alone points at a parser. A spike on every slot at once points at STT, context or latency.
4. Turn-level state progress. Snapshot `(node or task, filled slots)` after every agent turn. If the snapshot doesn't change for three agent turns in a row, the call is stalled, whatever the agent is saying. This catches loops the other three miss: tool-failure retries with varied wording, and flows stuck on a node that isn't collecting a slot.
One gotcha matters here. Count asks from the transcript, not from tool calls. If STT returns nothing (cause 1), the LLM never calls the tool, so a counter inside the tool stays at zero while the agent asks four times. The call logging guide lists the per-turn fields that make all four signals cheap to compute.
Starting thresholds
These are starting points to calibrate against labeled calls, not industry standards.
| Signal | Watch | Flag as loop | Notes |
|---|---|---|---|
| Consecutive agent-turn similarity | ≥ 0.70 | ≥ 0.80 on 2+ pairs | Ignore short fixed phrases like "Got it." |
| Asks per slot per call | 2 | 3 | The third ask is where recovery odds are worst |
| Agent turns with no state change | 2 | 3 | Exclude small-talk and hold states |
| Fleet re-ask rate, per slot | Week-over-week rise | Doubling after a deploy | Compare against the pre-change baseline |
| Duplicate user turns after a gap | Gap ≥ 2 s | 2+ per call | Points at cause 8 |
A repetition-loop detector in Python
Simplified, standard-library Python over a list of calls in JSON. Each agent turn needs `asks_slot`, `state` and `filled` fields, which your flow or task layer already knows.
# detect_loops.py - simplified. Usage: python3 detect_loops.py calls.json
import json, re, sys
from collections import defaultdict
from difflib import SequenceMatcher
SIM_FLAG = 0.80 # near-duplicate agent turns
REASK_LIMIT = 3 # asks of one slot before it counts as a loop
STALL_LIMIT = 3 # agent turns with no new slot filled and no state change
def norm(text):
text = re.sub(r"[^a-z0-9' ]+", " ", text.lower())
return re.sub(r"\s+", " ", text).strip()
def trigrams(text):
w = norm(text).split()
return {tuple(w[i:i + 3]) for i in range(len(w) - 2)} or {tuple(w)}
def similarity(a, b):
ta, tb = trigrams(a), trigrams(b)
jacc = len(ta & tb) / len(ta | tb)
seq = SequenceMatcher(None, norm(a), norm(b)).ratio()
return max(jacc, seq)
def analyze(call):
agent = [t for t in call["turns"] if t["role"] == "agent"]
flags, asks = [], defaultdict(int)
# Signal 1: near-duplicate consecutive agent turns
for prev, cur in zip(agent, agent[1:]):
s = similarity(prev["text"], cur["text"])
if s >= SIM_FLAG:
flags.append({"type": "near_duplicate", "turn": cur["idx"], "sim": round(s, 2)})
# Signal 2: same-question counter per slot
for t in agent:
if t.get("asks_slot"):
asks[t["asks_slot"]] += 1
if asks[t["asks_slot"]] == REASK_LIMIT:
flags.append({"type": "slot_reask", "turn": t["idx"], "slot": t["asks_slot"]})
# Signal 4: state-progress check
stall, last = 0, None
for t in agent:
snap = (t.get("state"), tuple(sorted((t.get("filled") or {}).items())))
stall = stall + 1 if snap == last else 0
last = snap
if stall == STALL_LIMIT:
flags.append({"type": "no_progress", "turn": t["idx"], "state": t.get("state")})
# Signal 3: re-ask rate for this call
requested = len(asks)
reasks = sum(n - 1 for n in asks.values())
return {
"call_id": call["call_id"],
"reask_rate": round(reasks / requested, 2) if requested else 0.0,
"max_asks_one_slot": max(asks.values(), default=0),
"loop": any(f["type"] in ("slot_reask", "no_progress") for f in flags),
"flags": flags,
}
if __name__ == "__main__":
calls = json.load(open(sys.argv[1]))
for c in calls:
for i, t in enumerate(c["turns"]):
t["idx"] = i
print(json.dumps(analyze(c)))Input shape, one call:
{"call_id": "c-1042", "turns": [
{"role": "agent", "text": "Thanks for calling. Can I get your date of birth?", "asks_slot": "dob", "state": "verify", "filled": {}},
{"role": "user", "text": "March third, eighty two"},
{"role": "agent", "text": "Sorry, I didn't get that. Can I get your date of birth?", "asks_slot": "dob", "state": "verify", "filled": {}},
{"role": "user", "text": "March. Third. Nineteen eighty two."},
{"role": "agent", "text": "Sorry, I didn't get that. Could I get your date of birth?", "asks_slot": "dob", "state": "verify", "filled": {}},
{"role": "user", "text": "I just told you"},
{"role": "agent", "text": "I'm sorry. Can I get your date of birth please?", "asks_slot": "dob", "state": "verify", "filled": {}}
]}Output for that call:
{"call_id": "c-1042", "reask_rate": 3.0, "max_asks_one_slot": 4, "loop": true,
"flags": [{"type": "near_duplicate", "turn": 4, "sim": 0.94},
{"type": "slot_reask", "turn": 4, "slot": "dob"},
{"type": "no_progress", "turn": 6, "state": "verify"}]}Look at what each signal caught. Similarity flagged one pair out of three re-asks. The slot counter fired on the third ask. The progress check fired at the fourth agent turn. On a clean call that fills `dob` and `phone` and reaches `confirm`, the detector returns no flags.
To extend it, add a duplicate-user-turn check for cause 8, then join each flag to the tool output and interrupted marker on the same turn. That join turns "loop detected" into "loop caused by a validator."
Fixes per cause
Every cause needs a cap and an exit. The cap stops the loop. The exit gets the caller what they called for.
| Cause | Root fix | Loop guard |
|---|---|---|
| 1. STT miss | Semantic end-of-turn for entity slots, keyterms, noise handling | After 2 misses, switch input mode |
| 2. Validator | Normalize spoken forms; log raw and parsed values | Read back the parsed value and confirm |
| 3. Tool failure | Return typed errors; say the system is down | Never re-ask for an argument that already failed twice |
| 4. Flow stuck | Handle every handler branch; precise completion-tool descriptions | Per-node turn budget, then force a transition |
| 5. Context loss | Slot values in structured state; re-inject on handoff | Check state before asking any slot |
| 6. Truncation | Fix false interruptions; repeat only the unheard part | Detect `interrupted=True` and don't count it as an ask |
| 7. Degeneration | Moderate temperature; collapse repeated turns in context | Similarity guard that rewrites the prompt |
| 8. Latency | Cut response latency; short acknowledgment while working | Drop a duplicate user turn that arrives mid-generation |
The escalation ladder
For any slot, use a fixed ladder. Each step changes the strategy, never just the wording. This follows the Bohus and Rudnicky result: changing the plan recovers more often than asking again.
1. Ask once, normally.
2. Re-ask with structure. Give an example or split the slot: "Let's do the month first."
3. Change the input mode. Spell-back ("I heard March 3, 1982. Is that right?"), letter-by-letter with a phonetic alphabet, or keypad entry. LiveKit ships a prebuilt `GetDtmfTask` for keypad or spoken digits. The DTMF testing guide covers carrier quirks.
4. Move on or hand off. Skip the slot if the task can proceed without it, verify another way, or offer a human with the collected context attached.

Pipecat Flows: count retries in state
Illustrative, using the consolidated-handler pattern from the Flows migration guide. `parse_spoken_date` and the node factories are yours.
from pipecat.flows import FlowArgs, FlowManager
MAX_DOB_TRIES = 2
async def record_dob(args: FlowArgs, flow_manager: FlowManager):
tries = flow_manager.state.get("dob_tries", 0) + 1
flow_manager.state["dob_tries"] = tries
dob = parse_spoken_date(args["spoken_dob"]) # handles "march third eighty two"
if dob:
flow_manager.state["dob"] = dob
return {"status": "ok", "dob": dob}, create_phone_node()
if tries >= MAX_DOB_TRIES:
return {"status": "switch_to_keypad"}, create_dob_keypad_node()
# Stay in this node, but tell the LLM to change approach
return {"status": "invalid", "next": "ask month, then day, then year"}, NoneLiveKit: cap the task, not just the prompt
Illustrative, using the documented `AgentTask`, `function_tool` and `complete()` API.
from livekit.agents import AgentTask, function_tool
class CollectDOB(AgentTask[str | None]):
def __init__(self, chat_ctx=None):
super().__init__(instructions="Collect the caller's date of birth.", chat_ctx=chat_ctx)
self._tries = 0
async def on_enter(self) -> None:
self.session.generate_reply(instructions="Ask for the caller's date of birth.")
@function_tool()
async def record_dob(self, spoken_dob: str) -> str:
"""Call this as soon as the caller says any date, exactly as spoken."""
self._tries += 1
dob = parse_spoken_date(spoken_dob)
if dob:
self.complete(dob)
return "saved"
if self._tries >= 2:
self.complete(None) # parent agent switches to keypad or a human
return "switching method"
return "Could not parse. Ask for the month, then the day, then the year."Both snippets share the gotcha from the detection section. They only count attempts the LLM sends to the tool. Pair them with a transcript-side counter so a slot the STT never captured still trips the ladder.
How to test for repetition loops
1. Instrument first. Add `asks_slot`, `state`, `filled`, `interrupted` and tool status to every agent turn. Without them, you're grepping transcripts by hand.
2. Run the detector on your last two weeks of calls. Rank slots by re-ask rate. That ranking is your test priority.
3. Build one scenario per cause. Use the matrix below. Each scenario should reproduce its cause on purpose, not wait for it to happen.
4. Run each scenario at least 10 times. LLM behavior varies run to run, so a single pass proves little. Report loop-call rate per scenario with the count, such as 3 of 10.
5. Pass criteria. No slot asked more than three times in any run. Zero stalls longer than three agent turns. Every scenario ends in either a filled slot or a handoff with context.
6. Re-run after every prompt, model, STT or flow change. Loops are regression-prone because each layer can reintroduce one. Synthetic callers make this cheap enough to do on every deploy.
Scenario matrix
| Cause | Scenario | How to reproduce |
|---|---|---|
| 1. STT miss | DOB with a 900 ms pause before the year | Scripted caller audio with an inserted pause |
| 1. STT miss | Phone number over street noise at 5 dB SNR on 8 kHz audio | Mix noise, downsample, play over SIP |
| 2. Validator | DOB in five spoken forms | "Three three eighty-two," "the third of March," and three more |
| 3. Tool failure | Lookup returns 500 twice | Fault-inject the tool endpoint |
| 4. Flow stuck | Partial match on account lookup | Return a result your handler doesn't branch on |
| 4. Flow stuck | "Wait, go back" mid-group | Correction phrase after task two of a `TaskGroup` |
| 5. Context loss | Name given before a handoff, needed after | Force a handoff at minute 5 |
| 6. Truncation | Cough 1 s into the DOB question | Overlay a cough on caller audio |
| 7. Degeneration | Three failed captures in a row, temperature 0 | Unintelligible caller audio for three turns |
| 8. Latency | Caller repeats the request after 2.5 s of silence | Add tool delay; script a repeat at 2.5 s |
Audio conditions matter as much as scripts. Run causes 1 and 6 over a real phone path, not a clean WebRTC link. The phone audio quality guide covers the codec and packet-loss conditions to include. If you'd rather have someone outside the team build and run this matrix, that's the job of an independent evaluation. It covers a pre-launch audit, a regression check after a model swap, or scoring production calls for re-ask rate.
Frequently asked questions
Why does my voice agent keep asking the same question?
Its state never advanced. The usual causes are an STT transcript that cut off or mangled the answer, a validator that rejects a spoken format, a tool error the LLM treats as bad input, a flow that didn't transition, or context that dropped the answer. Check the user turn and tool output just before each re-ask to see which one it is.
How is a repetition loop different from a handoff loop?
A repetition loop happens inside one agent, usually on one slot: the same question asked again and again. A handoff loop cycles the caller between agents, departments or queues. Both are livelocks, where activity continues without progress. Both need a counter plus a forced exit. Detect repetition loops per slot and handoff loops per transfer.
How many times should a voice agent re-ask before escalating?
Two asks, then change strategy. Research on spoken dialogue recovery found that asking a user to repeat had one of the lowest recovery rates. Moving on or giving help recovered far more often. A third identical ask has the worst odds. Use the third attempt for spell-back, keypad entry or a handoff.
Why does my LLM repeat the same response word for word?
Likelihood-based decoding at low temperature tends toward repetition, and research shows a self-reinforcement effect. Each repeat in the context makes the next repeat more likely. Raise temperature moderately, and collapse repeated agent turns in the chat history. Token frequency penalties mainly act on the response being generated, not on earlier turns.
Can I detect loops with embeddings alone?
Not reliably. Embedding similarity catches paraphrased repeats but can flag legitimate confirmations and misses stalls with varied wording. Combine it with a per-slot ask counter and a state-progress check. The counter catches re-asks whatever the wording. The progress check catches stalls where the agent says different things but nothing moves.
What is a good re-ask rate for a voice agent?
No public benchmark exists, so set your own baseline. Measure re-ask rate per slot over two weeks, then alert on increases after deploys. Entity slots such as dates of birth, phone numbers and IDs will sit higher than yes/no slots. A sudden rise across every slot at once usually points at STT, latency or context, not prompts.
Does a better STT model fix repetition loops?
Only cause 1, and only partly. Better STT reduces missed and mangled answers. It does nothing for validators that reject spoken formats, failed tools, stuck flows, dropped context or degenerate LLM output. Run the detector first and join flags to tool outputs and transcripts. Then you'll know whether STT is the cause before you pay to switch.
How do I stop my agent from re-asking after an interruption?
Make sure the barge-in was real, then repeat only the part the caller didn't hear. Fix false interruptions from coughs and backchannels first. Then don't count an interrupted agent turn as a completed ask in your retry counter. Otherwise one cough can push a caller up the escalation ladder.
The bottom line
A voice agent repeating itself is a symptom with eight causes across audio, code, orchestration, context, decoding and latency, and the fix depends on which layer broke. Count asks per slot from the transcript, check state progress every turn, cap retries at two, and make every retry change the strategy rather than the wording.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more