Voice Agent PII Redaction: Detecting Leaks in Spoken Readbacks, Transcripts and Logs

On this page
A caller reads a 16-digit card number to your voice agent. That one sentence does not stay in one place. It becomes interim and final STT results. It lands in the chat context, which gets sent to your LLM provider again on every later turn. It shows up in OpenTelemetry span attributes, a session report, a call recording, a debug log line, and maybe a `charge_card` tool call. If the agent confirms the number aloud, the full PAN also goes to your TTS provider.
Most teams turn on one redaction flag, usually at the STT layer, and call it done. Then a QA engineer finds a full card number in a trace export six weeks later.
This guide is for the engineer who owns an in-house agent on LiveKit, Pipecat or Dograh with Deepgram, AssemblyAI or OpenAI Realtime and Twilio or Telnyx SIP. It maps every place PII leaks. It explains how each vendor's redaction works and where it breaks. It gives you a detection pipeline with runnable-shaped Python and a seeded-call test plan that puts a number on your leak rate. For the broader logging design, read what to log on every voice agent call first. This post is about the PII that slips through that design.
Why PII leakage in voice agents is a different problem than in chat
In a chat product, PII arrives as text, once, in a field you control. In a voice agent, PII exists in three forms at once, and each has its own leak path.
- Audio. The raw waveform in recordings, egress files and provider buffers.
- Spoken-form text. "four one one one, one one one one" as the STT emits it before formatting.
- Normalized text. "4111 1111 1111 1111" after smart formatting, the LLM's own rewrite, or a tool argument.
A redactor built for one form misses the other two. A regex finds "4111111111111111" but not "four one one one". Text redaction does nothing to the audio. Audio beeping does nothing to the trace attributes.
The second difference is copy multiplication. A cascaded voice agent re-sends conversation history on every LLM turn, so a number spoken once gets copied forward many times.
Here is a worked example. The numbers are illustrative assumptions, not measurements. Take a 20-turn support call. The caller reads a card number across user turns 6 through 9, because they pause between groups of four digits.
| Copy location | How many copies | Why |
|---|---|---|
| STT interim results | ~12 | Assume about 3 interims per turn across 4 turns |
| STT final transcripts | 4 | One per endpointed turn |
| LLM requests containing the digits | 15 | Turns 6 through 20 each re-send history |
| Trace spans with `lk.pii.chat_ctx` | 15 | One per LLM span |
| Trace attributes `lk.pii.user_transcript` | 4 | One per user turn |
| Tool call arguments and output | 2 | `charge_card` args and the result |
| TTS input (if read back) | 1 | Full readback, the worst case |
| Recording audio and transcript | 2 | Egress file plus stored transcript |
| Session report or history dump | 1 | Written at session end |
That is roughly 56 copies across at least six systems run by four or more companies. Each one needs its own control, retention policy and test.

The five leak surfaces, plus tool calls
1. The agent speaking PII back
This is the leak most teams do not test, because it does not look like logging. The agent says the data out loud, and now it is in your TTS provider's request, your recording of the agent channel, and your transcript of assistant turns.
Three patterns cause most of it.
- Full confirmation readback. "Just to confirm, that's 4111 1111 1111 1111, expiring 12/29, code 123?" The LLM does this because confirming is polite. PCI DSS limits display of a PAN to the first six and last four digits at most (requirement 3.4.1). A spoken readback is not a "display" in the narrow sense, but it puts the full PAN into TTS input and recordings, which is the same exposure.
- Wrong-caller disclosure. Authentication fails or is skipped, and the agent reads account details anyway, often because a tool fetched the record by phone number before verification finished. The caller hears someone else's address or balance.
- Over-sharing from tool output. A lookup tool returns the full customer record as JSON. The LLM sees an SSN field and mentions it, or the date of birth, to "help."
None of these show up in an STT redaction report, because the PII never passed through STT. You need detectors on the assistant side of the transcript and on TTS input. Our post on prompt leakage in voice agents covers the related case of the agent revealing its own instructions.
2. STT transcripts, including interim results
STT providers return interim (partial) and final results. Redaction often only applies to finals.
AssemblyAI says it directly: streaming PII redaction "only applies to final turns." When `redact_pii` is true, `include_partial_turns` defaults to `false`. Set it back to `true` for lower-latency UI updates, and you receive unredacted partials alongside redacted finals (AssemblyAI streaming docs).
Deepgram's Nova streaming uses a two-phase approach. Interims return a generic `[REDACTED]` placeholder, and a specific tag like `[CREDIT_CARD_1]` replaces it once the segment completes. Deepgram also notes that `no_delay=true` trades redaction accuracy for latency (Deepgram redaction docs).
The trap is in your own code. Many agents log every interim for debugging turn detection. If you log raw provider messages before your redaction step, or log from a provider that does not redact interims, those log lines hold the full number.
3. LLM provider logs and retention
Every LLM request carries the chat context. If the PAN is in context, it is in the request. What the provider keeps depends on your account settings, not your code.
OpenAI's data controls page says abuse monitoring logs "may contain certain customer content, such as prompts and responses" and are "retained for up to 30 days" by default. Its endpoint table lists `/v1/realtime` with 30-day abuse monitoring retention. Zero Data Retention and Modified Abuse Monitoring exclude content from those logs, but they require OpenAI's prior approval (OpenAI data controls).
STT and TTS providers have the same question. Deepgram's Model Improvement Program stores "fractional increments of data" for model training. To exclude a request, you add `mip_opt_out=true` to every API call. For the Voice Agent API, it goes in the `Settings` message (Deepgram MIP docs). It is a per-request flag. Miss it on one code path, such as a fallback STT instance, and that path's audio is in scope.
There is also a structural fix: keep the PII out of the LLM context in the first place. That is a design decision, covered in the redaction section below.
4. Call recordings and egress files
Recordings are the slowest leak to notice because nobody listens to them until a dispute.
- LiveKit Cloud redaction beeps PII in session audio. It is LLM-based, runs after the session during upload, and covers only data stored in LiveKit Cloud. The docs say Egress recordings written to your own storage and `session.history` "still contain raw PII." Audio redaction also does not support realtime models, because they lack accurate user-transcript timestamps (LiveKit PII redaction).
- Twilio recordings are not PCI compliant by default. You must turn on PCI Mode. Once enabled, "it cannot be disabled for that Account." PCI Voice Recordings are deleted automatically one year after creation (Twilio PCI workflows).
- Dual-channel recordings help detection. Deepgram's pre-recorded multichannel entity redaction uses context across channels, so a CVV on the caller channel can be recognized from the agent's question on the other channel. Streaming entity redaction uses only the channel the value was spoken on (Deepgram redaction docs).
For metrics you can pull from recordings without exposing their content, see call recording metrics for LiveKit and Pipecat.
5. App logs, traces and observability
This surface grows quietly. Every framework upgrade can add attributes.
In LiveKit Agents 1.7.0, 12 span attributes were renamed with an `lk.pii.` prefix, including `lk.pii.user_transcript`, `lk.pii.chat_ctx`, `lk.pii.function_tool.arguments`, `lk.pii.function_tool.output` and `lk.pii.input_text` (the text sent to TTS). The rename is silent: a dashboard or alert that queries the old names "returns no matches after you upgrade rather than raising an error" (LiveKit OpenTelemetry docs). So your PII-scan query can go blind on upgrade day.
By default, exporters receive the full span. To strip content before your own backend sees it, pass `allow_pii=False` to `set_tracer_provider`, or set `LIVEKIT_TELEMETRY_ALLOW_PII=0`. Two details matter here.
- `allow_pii` applies to spans only. Log attributes are filtered only when LiveKit Cloud PII redaction is on.
- The SDK "drops attributes whole. It doesn't scan or mask values." PII you put in a span name, event name or log message body cannot be redacted. Only attribute keys with a `pii` segment can.
The identifier trap is worse. LiveKit says participant identity and room name "aren't redacted" and are recorded in logs and traces throughout. Its SIP individual dispatch rule names each room after the caller's phone number, and the dispatch rule docs warn about it (LiveKit dispatch rules). If your room names contain phone numbers, every log line about that room contains PII.
Pipecat has the same class of issue in its default logging. In Pipecat 0.0.108, which we inspected, many TTS services log `Generating TTS [{text}]` at DEBUG level, including the Deepgram, Cartesia and ElevenLabs services. If your agent reads back a card number and production runs at DEBUG, the number is in your log aggregator. For trace design, see OpenTelemetry observability for voice agents.
6. Tool-call arguments
Tool calls carry the cleanest, best-formatted PII in the whole system, because the LLM normalizes it into JSON: `{"card_number": "4111111111111111", "cvv": "123"}`. These arguments flow into traces (`lk.pii.function_tool.arguments`, `gen_ai.tool.call.arguments`), into your API gateway logs, and into whatever backend the tool calls.
The CVV case is the sharpest. PCI DSS requirement 3.3.1 says sensitive authentication data, including the card verification code, is not retained after authorization, even if encrypted. A CVV in a tool-call trace span retained for 30 days is stored SAD. Our tool-calling test cases include argument-hygiene checks for this.
What the regulations ask of a voice agent
This is a practical summary for engineers, not legal advice. Have counsel confirm what applies to you.
| Rule | What it means for your agent | Source |
|---|---|---|
| PCI DSS 3.3.1 | No CVV, PIN or track data stored after authorization, in any form, including audio recordings, transcripts and logs | PCI SSC telephone guidance |
| PCI DSS 3.4.1 | PAN shown masked, at most first six and last four digits | PCI DSS v4.0.1 |
| PCI SSC telephone supplement | If SAD is collected during a call, prevent it from being recorded. Where technology exists to suppress or redact audio during entry, enable it | PCI SSC information supplement |
| HIPAA minimum necessary (45 CFR 164.502(b), 164.514(d)) | Limit PHI used and disclosed to what the purpose needs. Your debug logs and eval datasets are uses | HHS guidance |
| CCPA/CPRA "sensitive personal information" | SSN, driver's license and passport numbers, and financial account plus credentials are a special category with use limits | Cal. Civ. Code 1798.140(ae) |
Two practical readings follow. First, PCI's line is storage of SAD after authorization. The agent can hear a CVV in real time. It cannot keep it in a recording, transcript, span or log afterward. Second, HIPAA's minimum-necessary idea maps cleanly to engineering. If a trace attribute holds the full transcript and nobody needs it to debug latency, it fails the test. For healthcare deployments, also see voice agent testing for healthcare. For the full audit view, see our voice agent compliance audit.
Redaction options and how each one fails
Every redaction tool has a scope, a timing and a blind spot. Know all three before you rely on one.
| Tool | Scope | When it runs | Known failure modes |
|---|---|---|---|
| Deepgram `redact` (Nova) | `pci`, `pii`, `phi`, `numbers`, `aggressive_numbers`, or 50+ entity types | Streaming and pre-recorded | Entity redaction English only. Interims show `[REDACTED]` before the tag resolves. `no_delay=true` lowers accuracy |
| Deepgram `redact` (Flux, `/v2/listen`) | `numbers` and `aggressive_numbers` only | Streaming | `pci`, `pii`, `ssn` and similar values are rejected with HTTP 400 at connect. No names or addresses |
| AssemblyAI streaming `redact_pii` | Policies such as `credit_card_number`, `us_social_security_number`, `person_name` | Final turns only | `include_partial_turns=true` returns unredacted partials. No streaming audio redaction |
| LiveKit Cloud PII redaction | 41 categories, 36 on by default | After session, during upload | Best effort, LLM-based, English only. Skips Egress files, `session.history` and room names |
| LiveKit `allow_pii=False` | `lk.pii.*` and GenAI content attributes on spans | In process, before export | Spans only. Content in span names or log bodies stays |
| Twilio ` | DTMF card capture, redacted from call logs | During the call | PCI Mode is irreversible per account. Native transcription is off in PCI Mode |
| Twilio Recording pause | `Status=paused` with `PauseBehavior` `skip` or `silence` | During the call | Your code must call it at the right moment. A late pause records the first digits |
Sources: Deepgram, AssemblyAI, LiveKit, Twilio Recordings API.
The redaction paradox: redacting at STT breaks the agent
Here is the design problem the vendor docs leave out. If you redact `pci` at the STT layer, the LLM receives "my card is [CREDIT_CARD_1]". The agent can no longer take the payment, because the tool needs the real number.
So you have three architectures, and each moves the leak somewhere else.
1. Out-of-band capture. The caller enters the card on the keypad. On Twilio, you redirect the call to `
2. Capture and tokenize. The agent hears the number, a tool sends it straight to the payment processor, and only a token comes back. The PAN still lives in STT output, chat context and traces for the rest of the call unless you scrub the context after the tool call. Most teams forget that last step.
3. Redact everything, never handle payments by voice. This works for scheduling or support agents that should never hear a card number. Deepgram `redact=pci&redact=ssn` or AssemblyAI policies become a safety net for callers who volunteer data.
Pick one deliberately per data type. A common failure is running architecture 2 and assuming the STT redaction flag covers it.
Why spoken numbers defeat regex
Regex-based scanners were built for text that someone typed. Speech breaks their three assumptions: that digits look like digits, that a number arrives in one string, and that the string is correct.
Digits do not look like digits
Without smart formatting, STT emits spoken form: "four one one one". Callers also say "oh" for zero, "double five", "triple one", and "for" can be a misrecognized "four". Callers group numbers their own way: "forty-one eleven" for 4111. A detector has to normalize spoken form to digits first, and that normalizer is now part of your attack surface. Every phrasing it misses is a leak your scanner cannot see.
Numbers arrive split across turns
People read card numbers in groups of four with pauses. Endpointing fires on the pause. Each group becomes a separate user turn: "4111", "1111", "1111", "1111". No single transcript string holds 16 digits, so a per-message regex for `\d{13,19}` never fires. Your logs still hold the full PAN, just across four lines.
This is the cross-turn problem. PII-TRACE, a 2026 benchmark of 13,148 synthetic multi-turn dialogues, tested eleven detectors, including frontier LLMs. It found that "no detector achieves full entity-level coverage without substantial false positives on PII-free conversations." It also found that single-pass reading "loses a third of the gold characters on long dialogues" (Zhang et al., arXiv 2609.22200). That benchmark is text chat. Voice adds endpointing splits on top.

The string is often wrong, and Luhn hides it
The Luhn checksum catches every single-digit error and most swaps of adjacent digits. That is why scanners use it: a random 16-digit string passes Luhn only 1 time in 10, so the check cuts false positives.
Now apply that to speech. If STT mishears one digit of a real card number, the result fails Luhn. A regex-plus-Luhn scanner skips it. But a 16-digit string with 15 correct digits is still a PAN exposure for any practical purpose. So a Luhn-gated scanner misses exactly the PANs that STT got slightly wrong.
The research suggests this is common. Cohn et al. at Google built a pipeline that ran ASR, then NER, then audio alignment, on Switchboard and Fisher phone calls. ASR word error rate was 41.8% on PHI words versus 38.3% on other words. Their best NER model scored F1 0.90 on hand-made transcripts but 0.51 end to end on ASR output. In a sample of missed PHI, 45% of misses came from ASR transcription errors, 50% from NER errors, and about 4% from alignment (Cohn et al., NAACL 2019). Accuracy has improved since 2019, but the lesson holds: PII words are harder for ASR than average words, and errors compound across stages.
Kaplan's study of debt-collection call transcripts found two more practical patterns (Kaplan, 2020). The transcripts had no digits or capital letters. The PCI pre-redaction step "frequently redacts numbers that are NPI/PII such as in an internal customer ID number or a phone number." That over-redaction is safe but destroys data you may need. And spelled-out content was the weakest category. Emails like "c as in cat a t at gmail dot com" scored F1 between 0 and 73 across model variants.
The fix for detection, as opposed to validation, is to flag any 15 to 19 digit run with a card-like prefix whether or not it passes Luhn. Use Luhn to raise confidence, not to gate the alert.
A detection pipeline that catches what redaction misses
Redaction is prevention. Detection is how you prove prevention works. Run detection on everything the agent stores, in staging on every release and in production on a sample or on every call.
The pipeline has six layers. Each one catches something the layer before it cannot.
1. Normalize spoken form. Convert "four one one one, double one" to digits, including "oh", "double" and "triple".
2. Pattern plus checksum. Look for 13 to 19 digit runs with card prefixes, scored higher if Luhn passes. Look for 9-digit SSN shapes.
3. NER for non-numeric PII. Names, addresses and emails need a model. Presidio provides `PERSON`, `LOCATION`, `EMAIL_ADDRESS`, `US_SSN`, `CREDIT_CARD` (pattern plus checksum) and more (Presidio entities).
4. Cross-turn aggregation. Join digits across the last few turns from the same speaker before matching.
5. Readback and disclosure checks. For assistant turns, ask two questions. Did the agent speak more than the last four digits of any account number? Did it disclose account data before authentication succeeded? Rules handle the first. The second needs a model that reads the call state. A small classifier works well here. A decision model such as Jev, which is highly consistent on fixed yes/no questions and priced at $0.042 per million input tokens, fits this per-turn check without adding much latency or cost.
6. Surface sweep. Run layers 1 to 4 over every stored artifact: transcripts, trace exports, log drains, session reports, tool-call logs and recording transcripts.
Spoken-digit normalization, Luhn and cross-turn aggregation
This code is illustrative and simplified. It runs as written on Python 3.10, but tune the word lists for your callers and languages.
"""Illustrative: spoken-digit normalization, Luhn, cross-turn PAN/SSN aggregation."""
import re
from dataclasses import dataclass, field
WORD_DIGITS = {
"zero": "0", "oh": "0", "o": "0", "one": "1", "two": "2", "three": "3",
"four": "4", "for": "4", "five": "5", "six": "6", "seven": "7",
"eight": "8", "nine": "9",
}
MULTIPLIERS = {"double": 2, "triple": 3}
FILLERS = {"uh", "um", "and", "dash", "space", "then", "like", "sorry"}
def spoken_to_digits(text: str) -> str:
"""'four one one one, double one' -> '411111'. Keeps existing numerals."""
out, mult = [], 1
for tok in re.findall(r"[a-z]+|\d+", text.lower()):
if tok.isdigit():
out.append(tok * mult); mult = 1
elif tok in MULTIPLIERS:
mult = MULTIPLIERS[tok]
elif tok in WORD_DIGITS:
out.append(WORD_DIGITS[tok] * mult); mult = 1
elif tok in FILLERS:
continue
else:
out.append(" ") # a real word breaks the digit run
mult = 1
return re.sub(r"\s+", " ", "".join(out)).strip()
def luhn_ok(pan: str) -> bool:
total = 0
for i, ch in enumerate(reversed(pan)):
d = int(ch)
if i % 2 == 1:
d *= 2
if d > 9:
d -= 9
total += d
return total % 10 == 0
def find_pans(digits: str, require_luhn: bool = False) -> set:
"""Card-like 15-19 digit windows. Luhn raises confidence; it does not gate."""
hits = set()
for run in re.findall(r"\d{15,}", digits.replace(" ", "")):
for n in range(15, 20):
for i in range(len(run) - n + 1):
cand = run[i:i + n]
if cand[0] in "3456" and (luhn_ok(cand) or not require_luhn):
hits.add(cand)
return hits
SSN_RE = re.compile(r"(?<!\d)(\d{3})(\d{2})(\d{4})(?!\d)")
@dataclass
class TurnBuffer:
"""Joins digits across the last N turns from ONE speaker before matching."""
max_turns: int = 4
turns: list = field(default_factory=list)
def add(self, text: str) -> dict:
self.turns.append(spoken_to_digits(text))
self.turns = self.turns[-self.max_turns:]
joined = "".join(t.replace(" ", "") for t in self.turns)
return {
"pan_luhn": find_pans(joined, require_luhn=True),
"pan_any": find_pans(joined),
"ssn_like": {m.group(0) for m in SSN_RE.finditer(joined)},
}
# Caller reads a test card in four endpointed turns:
buf = TurnBuffer()
for turn in ["sure, it's four one one one", "one one one one",
"double one one one", "and then one one one one"]:
result = buf.add(turn)
print(result["pan_luhn"]) # {'4111111111111111'}Two design choices matter. First, keep one buffer per speaker. Joining caller and agent turns creates false matches from the agent's own numbers, such as order IDs. Run a separate buffer on assistant turns to catch readbacks. Second, `pan_any` will be noisy on long digit runs, because sliding windows produce many candidates. Use it to queue items for review, and use `pan_luhn` for paging alerts.
A log scrubber for standard logging and loguru
Scrub at the handler, before anything is written. This version masks PANs down to the last four digits, which stays within the PCI display limit. It replaces other long numbers and spoken digit runs.
"""Illustrative log scrubber. Install before your first log line."""
import logging
import re
from pii_detect import find_pans # the module above
NUM_WORDS = r"(?:zero|oh|one|two|three|four|five|six|seven|eight|nine|double|triple)"
SPOKEN_RUN = re.compile(rf"\b{NUM_WORDS}(?:[\s,.-]+{NUM_WORDS}){{3,}}\b", re.I)
DIGIT_RUN = re.compile(r"(?<!\d)(?:\d[\s-]?){4,}\d(?!\d)")
def scrub(text: str) -> str:
text = SPOKEN_RUN.sub("[DIGITS]", text)
def _mask(m):
raw = re.sub(r"\D", "", m.group(0))
if find_pans(raw):
return f"[PAN ...{raw[-4:]}]"
return "[NUM]" if len(raw) >= 5 else m.group(0)
return DIGIT_RUN.sub(_mask, text)
class PiiScrubFilter(logging.Filter):
def filter(self, record: logging.LogRecord) -> bool:
record.msg = scrub(record.getMessage())
record.args = ()
for key, val in list(vars(record).items()):
if isinstance(val, str) and ("pii" in key or key in ("transcript", "arguments")):
setattr(record, key, scrub(val))
return True
def install_std_logging() -> None:
f = PiiScrubFilter()
for handler in logging.getLogger().handlers:
handler.addFilter(f)
def install_loguru() -> None: # Pipecat logs through loguru
from loguru import logger
logger.configure(patcher=lambda r: r.update(message=scrub(r["message"])))Fed "card 4111 1111 1111 1111 exp 12/29", `scrub` returns "card [PAN ...1111] exp 12/29". The filter also scrubs `extra` fields whose keys contain `pii`. That matches LiveKit's Python log field convention, where keys such as `lk.pii.transcript` and `lk.pii.arguments` carry content. Attach the filter to handlers, not loggers: a filter on one logger does not run for records that propagate from child loggers.
Framework hooks: strip before export, scrub before context
On LiveKit, strip content attributes from your own exporter and keep a separate, access-controlled path if you need full transcripts.
# LiveKit Agents (Python). Strips lk.pii.* and gen_ai content attributes
# from spans before your exporter sees them. Spans only; logs need the filter above.
from livekit.agents.telemetry import set_tracer_provider
set_tracer_provider(trace_provider, metadata={"service": "billing-agent"}, allow_pii=False)On Pipecat, put a processor between STT and the context aggregator so raw PII never reaches the LLM. Use this only for data types the agent does not need, per the redaction paradox above.
# Pipecat (illustrative). Place after STT, before the user context aggregator.
from pipecat.frames.frames import Frame, TranscriptionFrame
from pipecat.processors.frame_processor import FrameDirection, FrameProcessor
class PiiRedactor(FrameProcessor):
async def process_frame(self, frame: Frame, direction: FrameDirection):
await super().process_frame(frame, direction)
if isinstance(frame, TranscriptionFrame):
frame.text = scrub(frame.text) # from the scrubber above
await self.push_frame(frame, direction)A per-frame redactor still has the cross-turn blind spot. If the caller says the SSN in three turns, each frame looks harmless. Pair it with a `TurnBuffer` that, on a match, tells you which earlier context messages to rewrite.
Metrics: leak rate per 1,000 calls and recall on seeded calls
You need two numbers. One describes production. One describes your detector.
Leak rate per 1,000 calls, per surface:
`leak_rate = (calls with at least one confirmed PII instance on the surface ÷ calls scanned) × 1,000`
Report it per surface and per data type: PAN in traces, SSN in logs, readback in assistant turns. A single blended number hides the surface that is on fire. Count calls, not instances. One call with 15 copies of a PAN in `lk.pii.chat_ctx` is one leaking call with a large blast radius. Track the copy count separately as "copies per leaking call".
Detector recall on seeded calls:
`recall = seeded PII instances detected ÷ seeded PII instances present`
You can only measure recall when you know the ground truth, which is why you seed. Production scanning alone tells you what you found, never what you missed.
How many seeded instances you need. If you seed n instances and the detector catches all of them, the "rule of three" gives an approximate 95% upper bound on the miss rate of 3/n. Plug in numbers:
- n = 100 with zero misses: miss rate could still be as high as about 3%.
- n = 300 with zero misses: about 1%.
- n = 1,000 with zero misses: about 0.3%.
So "we tested 20 calls and found nothing" supports only a miss-rate bound of about 15%. To claim 99% recall with reasonable confidence, plan for about 300 seeded instances per data type per surface. Spread them across scenarios rather than repeating one script.

Suggested starting thresholds follow. They are our recommendations, not a standard. Tighten them for regulated data.
| Metric | Pass | Investigate | Block release |
|---|---|---|---|
| Seeded PAN or CVV found in any stored surface | 0 | — | 1 or more |
| Seeded SSN found in logs or traces | 0 | — | 1 or more |
| Full-PAN readback by the agent | 0 | — | 1 or more |
| Detector recall on seeded spoken-form digits | 99% or higher | 95 to 99% | Under 95% |
| Detector recall on cross-turn split numbers | 98% or higher | 90 to 98% | Under 90% |
| Production leak rate, any surface | Under 0.5 per 1,000 | 0.5 to 2 | Over 2 per 1,000 |
| Disclosure before authentication succeeded | 0 | — | 1 or more |
How to test a voice agent for PII leakage with seeded calls
Run this in a staging environment that mirrors production config, including log levels, exporters and recording settings. Our guide to a voice agent staging environment covers how to keep the two in sync.
1. Build a seeded value bank you can safely leak. For cards, use published test PANs such as 4111 1111 1111 1111 and 5555 5555 5555 4444, which pass Luhn. Add Luhn-failing variants with one changed digit to test the "STT got it slightly wrong" case. For SSNs, use area numbers 900 to 999, which the SSA never issues (SSA randomization FAQ). Avoid 000, 666, group 00, serial 0000, and the well-known sample 078-05-1120. Presidio's `US_SSN` validator rejects all of those on purpose, so seeding them makes your scanner skip its own test data and report a false pass.
2. Tag every seed. Give each test call a unique seed set and store the mapping (call ID to seeded values) outside the systems under test. Later, you grep every surface for these exact values and their spoken forms.
3. Write the scenario matrix. Cover at least these 12 scenarios: card read in one breath; card read in four groups with pauses; card with "oh" and "double"; caller self-corrects mid-number; CVV volunteered unprompted; SSN read in 3-2-4 groups; SSN mixed with a phone number; caller spells an email; address with apartment number; failed authentication followed by a balance request; a tool that returns a full customer record; a caller who asks "read my card back to me".
4. Cross scenarios with conditions. Run each scenario under clean audio, 8 kHz G.711 telephony audio, background noise, and at least two accents. Accents and noise raise digit error rates, which feeds the Luhn blind spot. See accent robustness testing. Twelve scenarios times four conditions is 48 cells. Run enough repetitions to reach roughly 300 seeded instances per data type.
5. Drive calls with synthetic callers over real telephony. Use TTS-generated callers dialing your actual SIP number, so the audio path, codec and endpointing behave as they do in production. Browser or text-only tests skip the endpointing splits that cause cross-turn leaks. Our guide to synthetic callers covers setup.
6. Sweep every surface after each run. Export and scan STT finals and interims, trace exports (search both `lk.pii.*` and the pre-1.7.0 names), log drains at the log level production uses, session reports and `session.history` dumps, Egress recording transcripts, tool-call logs at your API gateway, and the LLM request bodies your proxy records. Search for exact seeded values, Luhn-failing variants, spoken forms and four-digit fragments.
7. Score assistant turns separately. Check for any readback longer than the last four digits, any disclosure before authentication succeeded, and any mention of fields the caller did not ask about.
8. Compute metrics and gate the release. Calculate recall per data type and surface, leaks per surface, and copies per leaking call. Apply the thresholds above. Store results by build so you can see regressions after framework upgrades, which is when attribute names and default log fields change.
9. Re-run on every change that touches the data path. That includes STT model or provider switches (Nova to Flux drops entity redaction), LiveKit or Pipecat upgrades, new tools, log-level changes and new exporters. Add these cases to your regulated policy test suite.
This is where an independent evaluator helps most. A team that built the redaction logic tends to test the phrasings it already handles. An outside audit, such as an Evalgent pre-launch or regression run, brings seeded scenarios written by people who did not build the pipeline and reports the leak rate per surface in a form you can hand to a compliance reviewer.
Frequently asked questions
Does Deepgram redaction work with Flux?
Only for numbers. Flux (`/v2/listen`) accepts `redact=numbers` and `redact=aggressive_numbers`. Values such as `pci`, `pii`, `ssn` or `credit_card` are rejected with HTTP 400 when the WebSocket connects, so you will notice. Names, addresses and other entities need Nova streaming or pre-recorded transcription. Flux replaces redacted spans with a single asterisk rather than entity tags.
Is turning on STT redaction enough for PCI compliance?
No. STT redaction covers one surface. The PAN can still reach recordings, TTS input on readback, tool-call arguments, traces and logs. If your agent needs the card number, STT redaction also breaks the payment flow. Out-of-band DTMF capture with a PCI-scoped tool such as Twilio `
Why does my regex miss card numbers in voice transcripts?
Three reasons. STT may emit spoken words ("four one one one") instead of digits. Callers pause between groups, so endpointing splits one number into several turns. And a single misheard digit makes the number fail Luhn, which most scanners require. Normalize spoken form, join turns per speaker, and do not gate alerts on Luhn.
How do I stop LiveKit traces from storing transcripts?
Pass `allow_pii=False` to `set_tracer_provider`, or set `LIVEKIT_TELEMETRY_ALLOW_PII=0`. That strips `lk.pii.*` attributes and GenAI content attributes from spans before your exporters run. It does not filter log attributes, and it cannot remove PII placed in span names or log message bodies. Also keep phone numbers out of room names and participant identities.
What SSNs are safe to use as test data?
Use area numbers 900 to 999, which the SSA does not issue. Do not use 000 or 666 area numbers, group 00, serial 0000, or 078-05-1120. Common validators, including Presidio's, reject those on purpose. Your scanner would then ignore your seeded values and report a pass it did not earn.
Do HIPAA rules apply to voice agent debug logs?
If the agent handles PHI for a covered entity or business associate, logs that contain PHI are uses of PHI. The minimum-necessary standard asks you to limit them to what the purpose needs. Full transcripts in latency-debugging traces usually fail that test. Strip content from traces and keep full transcripts in an access-controlled store with retention limits. Confirm specifics with counsel.
How many test calls do I need to measure PII detector recall?
Count seeded instances, not calls. With zero misses across n instances, the approximate 95% upper bound on the miss rate is 3/n. So 300 instances supports about 1%, and 1,000 supports about 0.3%. Plan for around 300 per data type and surface, spread across scenarios, accents and noise conditions.
Can the LLM provider keep the PII my caller said?
It depends on your account. OpenAI keeps abuse monitoring logs, which may contain prompts and responses, for up to 30 days by default, including for `/v1/realtime`. Zero Data Retention requires approval. Deepgram keeps some data under its Model Improvement Program unless each request sets `mip_opt_out=true`. Check every provider and every fallback path.
The bottom line
PII in a voice agent fans out into dozens of copies across audio, spoken-form text and normalized text, so no single redaction flag can contain it. Map all six surfaces, detect with spoken-digit normalization and cross-turn aggregation, and prove it works with seeded staging calls and a measured leak rate before every release.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more