Blog
Page 5 of 32

Voice agent observability tool: a 14-point buyer's and builder's checklist
What a voice agent observability tool must capture: per-turn traces, stereo audio, signal health, caller-heard latency, audit trails. Scored, testable.

Pipecat vs LiveKit (2026): Which to Choose, and How to Prove It
Pipecat vs LiveKit: LiveKit for phone, SIP and multi-party agents; Pipecat for complex Python pipelines. Verdict table, same agent in both, 9 deep dives.

AI voice agent testing: the complete guide to layers, failure modes and regression (2026)
AI voice agent testing in five layers: unit, audio replay, simulated phone calls, load and production scoring. Failure taxonomy, thresholds and regression.

Deepgram Testing for Voice Agents: Latency, Accuracy and Load (2026)
Deepgram testing for voice agents: measure latency and WER, load test concurrency limits, and get direct answers to 11 common Deepgram questions.

Voice agent regression testing: how to detect regressions from model updates you don't control
Voice agent regression testing for changes you don't control: pin each layer, run a nightly pinned-vs-latest canary, and use paired stats to detect drops.

Jev Accuracy in Voice Agent Evals: Where It Fails on Calls
Jev scores clean calls well, but it can't abstain. See where Jev accuracy breaks on real calls, how forced answers skew QA metrics, and how to fix it.

Replacing Manual Test Calls on a LiveKit or Pipecat Agent: How to Automate Voice Agent Testing With a Small Team
Break a 30-minute manual test call session into 16 checks, then automate them in four layers on LiveKit or Pipecat, with cost math and flake control.

Voice Agent Prompt Versioning: Shipping Prompt and Flow Changes to an In-House Voice Agent Without Breaking Live Calls
Voice agent prompt versioning for LiveKit and Pipecat: pin each call to a release, gate on pass^k, size the canary, and roll back without dropping calls.

Deepgram Flux vs Nova-3 for Voice Agents: Accuracy, Turn Detection, Latency and Cost
Flux or Nova-3 for a phone voice agent? Verified features, pricing at 50k to 500k calls, what each latency number means, and a bake-off you can run.