Evalgent

Blog

Page 5 of 32

Voice agent observability tool: a 14-point buyer's and builder's checklist
Voice AI Evaluation
21 min read

Voice agent observability tool: a 14-point buyer's and builder's checklist

What a voice agent observability tool must capture: per-turn traces, stereo audio, signal health, caller-heard latency, audit trails. Scored, testable.

Updated October 5, 2026Read more
Pipecat vs LiveKit (2026): Which to Choose, and How to Prove It
Voice AI Testing
18 min read

Pipecat vs LiveKit (2026): Which to Choose, and How to Prove It

Pipecat vs LiveKit: LiveKit for phone, SIP and multi-party agents; Pipecat for complex Python pipelines. Verdict table, same agent in both, 9 deep dives.

Updated October 5, 2026Read more
AI voice agent testing: the complete guide to layers, failure modes and regression (2026)
Voice AI Testing
21 min read

AI voice agent testing: the complete guide to layers, failure modes and regression (2026)

AI voice agent testing in five layers: unit, audio replay, simulated phone calls, load and production scoring. Failure taxonomy, thresholds and regression.

Updated October 5, 2026Read more
Deepgram Testing for Voice Agents: Latency, Accuracy and Load (2026)
Testing Strategies
19 min read

Deepgram Testing for Voice Agents: Latency, Accuracy and Load (2026)

Deepgram testing for voice agents: measure latency and WER, load test concurrency limits, and get direct answers to 11 common Deepgram questions.

Updated October 5, 2026Read more
Voice agent regression testing: how to detect regressions from model updates you don't control
Voice AI Evaluation
16 min read

Voice agent regression testing: how to detect regressions from model updates you don't control

Voice agent regression testing for changes you don't control: pin each layer, run a nightly pinned-vs-latest canary, and use paired stats to detect drops.

Updated October 5, 2026Read more
Jev Accuracy in Voice Agent Evals: Where It Fails on Calls
Voice AI Evaluation
23 min read

Jev Accuracy in Voice Agent Evals: Where It Fails on Calls

Jev scores clean calls well, but it can't abstain. See where Jev accuracy breaks on real calls, how forced answers skew QA metrics, and how to fix it.

October 4, 2026Read more
Replacing Manual Test Calls on a LiveKit or Pipecat Agent: How to Automate Voice Agent Testing With a Small Team
Voice AI Testing
25 min read

Replacing Manual Test Calls on a LiveKit or Pipecat Agent: How to Automate Voice Agent Testing With a Small Team

Break a 30-minute manual test call session into 16 checks, then automate them in four layers on LiveKit or Pipecat, with cost math and flake control.

October 4, 2026Read more
Voice Agent Prompt Versioning: Shipping Prompt and Flow Changes to an In-House Voice Agent Without Breaking Live Calls
Voice AI Testing
25 min read

Voice Agent Prompt Versioning: Shipping Prompt and Flow Changes to an In-House Voice Agent Without Breaking Live Calls

Voice agent prompt versioning for LiveKit and Pipecat: pin each call to a release, gate on pass^k, size the canary, and roll back without dropping calls.

October 4, 2026Read more
Deepgram Flux vs Nova-3 for Voice Agents: Accuracy, Turn Detection, Latency and Cost
Voice AI Testing
24 min read

Deepgram Flux vs Nova-3 for Voice Agents: Accuracy, Turn Detection, Latency and Cost

Flux or Nova-3 for a phone voice agent? Verified features, pricing at 50k to 500k calls, what each latency number means, and a bake-off you can run.

October 4, 2026Read more