Evalgent

Blog

Page 4 of 32

How to Evaluate Voice Variations With A/B Tests in Deepgram (2026)
Voice AI Testing
24 min read

How to Evaluate Voice Variations With A/B Tests in Deepgram (2026)

Evaluate voice variations with A/B tests in Deepgram: pin Aura-2 or Flux TTS per session, randomize by caller, size the sample, and avoid peeking errors.

Updated October 5, 2026Read more
The Best Way to Load Test Deepgram's Voice Agent API (2026 Harness)
Voice AI Testing
24 min read

The Best Way to Load Test Deepgram's Voice Agent API (2026 Harness)

The best way to load test Deepgram's Voice Agent API: Erlang B sizing, Poisson arrivals, real-time 20 ms audio, p99 turn latency, and an asyncio harness.

Updated October 5, 2026Read more
GPT-Live API architecture: build voice agents with gpt-live-1
Voice AI Evaluation
19 min read

GPT-Live API architecture: build voice agents with gpt-live-1

GPT-Live API architecture explained: how gpt-live-1 splits voice from backend reasoning, plus WebRTC and SIP session code, delegation, and pricing.

Updated October 5, 2026Read more
Menu modifiers in voice ordering: how to test food ordering voice agents
Voice AI Testing
19 min read

Menu modifiers in voice ordering: how to test food ordering voice agents

What menu modifiers are, how POS menus model them, why voice agents get them wrong, and how to test drive-thru and phone ordering before rollout.

Updated October 5, 2026Read more
Grok Voice Agent Builder (xAI) Pricing: $0.08/min and Setup (Oct 2026)
Voice AI Evaluation
17 min read

Grok Voice Agent Builder (xAI) Pricing: $0.08/min and Setup (Oct 2026)

Grok Voice Agent Builder (xAI) pricing, October 2026: $0.08 per minute plus $0.01 telephony. Cost math, models, voices, cloning and phone setup.

Updated October 5, 2026Read more
How to Monitor Voice Agents in Production: Tool Failures, Latency Spikes and Dead Air
Voice AI Evaluation
19 min read

How to Monitor Voice Agents in Production: Tool Failures, Latency Spikes and Dead Air

How to monitor voice agents in production on LiveKit or Pipecat: tool-call failure rates, p95 latency per segment, dead-air detection and burn-rate alerts.

Updated October 5, 2026Read more
How to A/B Test Voice Agent Prompts: Offline Pairs, Live Splits, and the Sample-Size Math
Voice AI Testing
18 min read

How to A/B Test Voice Agent Prompts: Offline Pairs, Live Splits, and the Sample-Size Math

How to A/B test voice agent prompts on LiveKit or Pipecat: paired offline runs, caller-level live splits, blind judging, and sample-size math that holds.

Updated October 5, 2026Read more
Best LLM for voice agents (October 2026): latency, tool calling, cost per minute
Voice AI Evaluation
17 min read

Best LLM for voice agents (October 2026): latency, tool calling, cost per minute

Best LLM for voice agents, October 2026: GPT-6, Claude, Gemini, Grok and open models compared on first-token latency, tool calling and cost per minute.

Updated October 5, 2026Read more
Full-duplex voice agents: turn-based vs full-duplex, models, testing
Voice AI Evaluation
18 min read

Full-duplex voice agents: turn-based vs full-duplex, models, testing

Turn-based vs full-duplex voice agents: one waits for silence, the other listens while it speaks. Compare GPT-Live, Moshi and PersonaPlex, and how to test.

Updated October 5, 2026Read more