Evalgent

Blog

Page 3 of 6

Voice agent testing vs evaluation vs monitoring vs observability
Voice AI Evaluation
11 min read

Voice agent testing vs evaluation vs monitoring vs observability

Voice agent testing vs monitoring vs evaluation vs observability: what each term means, when to use it, and how they fit into one loop.

July 3, 2026Read more
OpenTelemetry observability for AI voice agents
Voice AI Evaluation
11 min read

OpenTelemetry observability for AI voice agents

OpenTelemetry for voice agents traces every turn across STT, LLM, and TTS. Learn the GenAI conventions, the spans to emit, and how to instrument it.

July 3, 2026Read more
How to improve WER: reducing word error rate for voice agents
Voice AI Evaluation
10 min read

How to improve WER: reducing word error rate for voice agents

Improve WER for your voice agent with custom vocabulary, cleaner audio, model choice, and tuning. A practical guide to reducing word error rate per cohort.

July 3, 2026Read more
TTS evaluation: how to test text-to-speech for voice agents
Voice AI Evaluation
10 min read

TTS evaluation: how to test text-to-speech for voice agents

TTS evaluation measures text-to-speech naturalness, intelligibility, latency, and pronunciation. Learn the metrics, MOS, and how to test a voice.

July 3, 2026Read more
STT evaluation: how to test speech-to-text for voice agents
Voice AI Evaluation
10 min read

STT evaluation: how to test speech-to-text for voice agents

STT evaluation measures speech-to-text accuracy under real conditions: accents, noise, and jargon. Learn the metrics, the method, and what benchmarks miss.

July 3, 2026Read more
Word error rate (WER): the complete guide for voice agents
Voice AI Evaluation
10 min read

Word error rate (WER): the complete guide for voice agents

Word error rate measures speech-to-text accuracy via substitutions, deletions, and insertions over total words. Learn how to calculate it and its limits.

July 3, 2026Read more
Best LLM for voice agents: how to choose
Voice AI Evaluation
10 min read

Best LLM for voice agents: how to choose

The best LLM for voice depends on latency, instruction following, and tool calling. Compare GPT, Claude, and Gemini, and learn how to choose.

July 3, 2026Read more
Self-improving voice agents: how feedback loops make agents better
Voice AI Evaluation
11 min read

Self-improving voice agents: how feedback loops make agents better

Self-improving voice agents get better from production feedback. Learn how the loop works, where it differs from retraining, and how to keep it safe.

July 3, 2026Read more
Full-duplex voice agents: how simultaneous speech changes voice AI
Voice AI Evaluation
10 min read

Full-duplex voice agents: how simultaneous speech changes voice AI

Full-duplex voice agents listen and speak at once, enabling barge-in and interruptions. Learn how they differ from half-duplex and how to test them.

July 3, 2026Read more