# Evalgent > Evalgent is the AI voice agent testing and evaluation platform. Run scenario-based testing, simulate real users with synthetic callers, measure custom metrics, and monitor production voice agents for reliability — voice agent QA built for teams shipping production voice agents at scale. ## Pages - [Homepage](https://www.evalgent.com/): The AI voice agent testing and evaluation platform built for production reliability — scenarios, synthetic callers, custom metrics, and monitoring. - [Platform](https://www.evalgent.com/platform): The five building blocks of voice agent testing: scenarios, caller profiles, metrics, evaluations, and reviews. Provider-agnostic release gate. - [Industries](https://www.evalgent.com/industries): Voice agent testing by industry — scenario coverage, caller profiles, and outcome verification tuned to each vertical's workflows and systems. - [Test AI voice agents](https://www.evalgent.com/test-ai-voice-agent): AI voice agent testing across accents, background noise, and edge cases before production deployment. - [Iterate AI voice agents](https://www.evalgent.com/iterate-ai-voice-agent): Regression testing on prompt changes, model updates, and flow improvements before release. - [Monitor AI voice agents](https://www.evalgent.com/monitor-ai-voice-agent): Continuous monitoring for production voice agents — drift detection and emerging-failure alerts (coming soon). - [Scenarios](https://www.evalgent.com/platform/scenarios): Define structured test scenarios with objectives and success criteria, manually or auto-generated from agent instructions. - [Profiles](https://www.evalgent.com/platform/profiles): Configure synthetic callers across accent, speech pace, interruptions, latency, and behaviour to stress-test under realistic conditions. - [Metrics](https://www.evalgent.com/platform/metrics): Define custom telemetry and LLM-judged metrics for tone, accuracy, latency, and task completion. - [Evaluations](https://www.evalgent.com/platform/evaluations): Run automated evaluation campaigns combining scenarios, profiles, and metrics with pass/fail verdicts and evidence. - [Reviews](https://www.evalgent.com/platform/reviews): Human-in-the-loop review — challenge LLM verdicts, correct outcomes, and keep a full audit trail. - [Contact](https://www.evalgent.com/contact): Reach the Evalgent team by email or phone, or book a 30-minute call. - [Blog](https://www.evalgent.com/blog): Practical guides on voice agent testing, synthetic callers, regression testing, production monitoring, and reliability. ## Industries - [Healthcare](https://www.evalgent.com/industries/healthcare): Booking, refills, triage, and insurance verification, with outcome checks against the EHR. - [Financial Services](https://www.evalgent.com/industries/financial-services): Payments, collections, fraud verification, and lending, verified against core banking systems. - [Insurance](https://www.evalgent.com/industries/insurance): FNOL claims intake, quotes, renewals, and billing with accurate data capture. - [E-commerce](https://www.evalgent.com/industries/ecommerce): Order status, returns, support, and retention for high-volume, multilingual shoppers. - [Hospitality](https://www.evalgent.com/industries/hospitality): Reservations, guest services, and loyalty in every language guests call in. - [Logistics](https://www.evalgent.com/industries/logistics): Dispatch, ETAs, and delivery recovery under real vehicle noise and clipped driver speech. - [Automotive](https://www.evalgent.com/industries/automotive): Service booking, recall outreach, and sales lead qualification for dealerships. - [Real Estate](https://www.evalgent.com/industries/real-estate): Lead qualification, viewing scheduling, listings, and tenant requests. - [Recruiting](https://www.evalgent.com/industries/recruiting): Fair, consistent candidate screening, interview scheduling, and onboarding. - [Education](https://www.evalgent.com/industries/education): Admissions, student support, and tuition through seasonal demand spikes. ## Guides - [How to test tool calling in AI voice agents](https://www.evalgent.com/resources/guides/test-tool-calling-voice-agents): Tool calling fails when a voice agent picks the wrong function, fills bad arguments, or acts silently. Learn how to test tool calling and catch it. ## Blog Posts - [Why AI voice agents fail in production (and how to prevent it)](https://www.evalgent.com/blog/why-voice-agents-fail-in-production): AI voice agents that ace demos still break in production. Learn the 5 root causes, how to test for each, and what production readiness actually means. - [LLM as judge for voice agents: the hidden limits of transcript evaluation](https://www.evalgent.com/blog/llm-as-judge-voice-ai-evaluation-limits): LLM-as-judge scores voice agents high while real failures slip through. The 5 blind spots of transcript scoring, and what outcome-based evaluation looks like. - [Conversational AI testing: the complete voice agent stress testing guide](https://www.evalgent.com/blog/stress-testing-voice-ai): Systematically stress-test voice agents to find breaking points across noise, accents, interruptions, and latency, before real users hit them. - [Voice agent regression testing: why LLM updates break production](https://www.evalgent.com/blog/llm-update-regression-voice-agent): LLM updates improve benchmarks but break voice agents in 5 predictable ways. How to detect and prevent regressions after every model or prompt change. - [How to automate voice agent testing: synthetic callers vs manual QA](https://www.evalgent.com/blog/synthetic-callers-for-voice-agent-testing): Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring. - [ElevenLabs voice agent testing guide: what to check before going live](https://www.evalgent.com/blog/elevenlabs-voice-agent-testing-guide): Test your ElevenLabs voice agent before launch: scenario gaps, real user behaviour, tool calls, concurrency limits, and voice-quality regression. - [Vapi voice agent testing guide: what to check before going live](https://www.evalgent.com/blog/vapi-voice-agent-testing-guide): Test your Vapi voice agent before going live. Covers BYOK costs, Squads handoff gaps, webhook failures, and prompt regression before real users find them. - [LiveKit voice agent testing guide: what to check before going live](https://www.evalgent.com/blog/livekit-voice-agent-testing-guide): LiveKit Agents testing guide: turn detection failures, SIP integration gaps, worker scaling, and prompt regression — what to check before going live. - [Deepgram STT testing guide for voice agents: what the benchmarks don't tell you](https://www.evalgent.com/blog/deepgram-stt-voice-agent-testing-guide): A practical guide to load-testing and monitoring Deepgram STT in voice agents: latency, WER under noise and accents, and error compounding before you go live. - [Pipecat voice agent testing guide: what to check before going live](https://www.evalgent.com/blog/pipecat-voice-agent-testing-guide): Before your Pipecat agent goes live, check VAD, frame drops, transport, the Flows state machine, and pipeline regression. A pre-launch testing checklist. - [Voice agent stack: the complete guide](https://www.evalgent.com/blog/voice-agent-stack-guide): The complete voice agent stack: STT, LLM, TTS, orchestration, and telephony. Latency budget, cost per layer, build vs buy, and what to test before production. - [What is a voice agent? The complete guide](https://www.evalgent.com/blog/what-is-a-voice-agent): What is a voice agent? How it works, how it differs from IVR and chatbots, use cases by industry, and what production-ready deployment actually requires. - [AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice](https://www.evalgent.com/blog/ai-agent-testing-vs-voice-agent-testing): AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss. - [You don't read AI-generated code. Why are you listening to every call?](https://www.evalgent.com/blog/eval-driven-development): Nobody reads AI-generated code anymore. Tests became mandatory. Voice agents are next — eval-driven development is how reliable voice AI gets shipped. - [The ROI of AI Voice Agent Testing Isn't What You Think It Is](https://www.evalgent.com/blog/roi-of-ai-voice-agent-testing): Cost-per-test is the wrong way to measure voice agent testing ROI. Six capabilities that change how teams ship voice AI — and what they're worth. - [Why Production Voice Agents Are Becoming Graphs, Not Prompts](https://www.evalgent.com/blog/graph-based-voice-agents): Graph-based voice agents externalise control flow that prompt-based agents trust the LLM to handle. Why this is the architecture winning production deployments. - [Should You Build a Graph-Based or Prompt-Based Voice Agent?](https://www.evalgent.com/blog/graph-based-vs-prompt-based-voice-agents): Graph-based vs prompt-based voice agents: a decision framework. Five questions, a clear matrix, and three hybrid patterns most teams actually ship. - [Testing Graph-Based Voice Agents: The Three-Tier Model](https://www.evalgent.com/blog/testing-graph-based-voice-agents): Testing graph-based voice agents requires three tiers — node, subgraph, end-to-end voice. Most teams stop at tier one. Here's what production-grade looks like. - [AI voice agent testing: the complete guide](https://www.evalgent.com/blog/ai-voice-agent-testing): AI voice agent testing checks your agent through real speech: accents, interruptions, latency, and edge cases. Learn what to test and which metrics matter. - [Vapi vs Retell: which voice agent platform should you choose?](https://www.evalgent.com/blog/vapi-vs-retell): Vapi vs Retell compared on architecture, pricing, latency, telephony, and scale. A clear decision framework for choosing a voice agent platform in 2026. - [LiveKit vs Vapi: which voice agent platform should you choose?](https://www.evalgent.com/blog/livekit-vs-vapi): LiveKit vs Vapi compared on architecture, pricing, latency, telephony, and scale. A clear framework for choosing a voice agent platform in 2026. - [Pipecat vs LiveKit: which voice agent framework should you choose?](https://www.evalgent.com/blog/pipecat-vs-livekit): Pipecat vs LiveKit compared on architecture, turn detection, latency, telephony, and pricing. A clear framework for choosing a voice agent stack in 2026. - [Voice agent observability: how to monitor voice agents in production](https://www.evalgent.com/blog/voice-agent-observability): Voice agent observability means monitoring live agents across audio, latency, and behaviour. Learn what to track and how it differs from testing. - [SOP-based voice agents: how standard operating procedures make agents reliable](https://www.evalgent.com/blog/sop-based-voice-agents): SOP-based voice agents follow standard operating procedures instead of relying on a prompt. Learn how SOPs improve reliability, control, and testability. - [Testing SOP-based voice agents: how to verify procedure adherence](https://www.evalgent.com/blog/testing-sop-based-voice-agents): Testing SOP-based voice agents means verifying every step and branch of the procedure. Learn SOP adherence testing, step coverage, and metrics. - [Full-duplex voice agents: how simultaneous speech changes voice AI](https://www.evalgent.com/blog/full-duplex-voice-agents): Full-duplex voice agents listen and speak at once, enabling barge-in and interruptions. Learn how they differ from half-duplex and how to test them. - [Self-improving voice agents: how feedback loops make agents better](https://www.evalgent.com/blog/self-improving-voice-agents): Self-improving voice agents get better from production feedback. Learn how the loop works, where it differs from retraining, and how to keep it safe. - [Best LLM for voice agents: how to choose](https://www.evalgent.com/blog/best-llm-for-voice-agents): The best LLM for voice depends on latency, instruction following, and tool calling. Compare GPT, Claude, and Gemini, and learn how to choose. - [Word error rate (WER): the complete guide for voice agents](https://www.evalgent.com/blog/word-error-rate-voice-agents): Word error rate measures speech-to-text accuracy via substitutions, deletions, and insertions over total words. Learn how to calculate it and its limits. - [STT evaluation: how to test speech-to-text for voice agents](https://www.evalgent.com/blog/stt-evaluation-voice-agents): STT evaluation measures speech-to-text accuracy under real conditions: accents, noise, and jargon. Learn the metrics, the method, and what benchmarks miss. - [TTS evaluation: how to test text-to-speech for voice agents](https://www.evalgent.com/blog/tts-evaluation-voice-agents): TTS evaluation measures text-to-speech naturalness, intelligibility, latency, and pronunciation. Learn the metrics, MOS, and how to test a voice. - [How to improve WER: reducing word error rate for voice agents](https://www.evalgent.com/blog/how-to-improve-wer-voice-agents): Improve WER for your voice agent with custom vocabulary, cleaner audio, model choice, and tuning. A practical guide to reducing word error rate per cohort. - [OpenTelemetry observability for AI voice agents](https://www.evalgent.com/blog/opentelemetry-observability-voice-agents): OpenTelemetry for voice agents traces every turn across STT, LLM, and TTS. Learn the GenAI conventions, the spans to emit, and how to instrument it. - [Voice agent testing vs evaluation vs monitoring vs observability](https://www.evalgent.com/blog/voice-agent-testing-vs-monitoring-vs-observability): Voice agent testing vs monitoring vs evaluation vs observability: what each term means, when to use it, and how they fit into one loop. - [A/B testing voice agents: how to compare versions the right way](https://www.evalgent.com/blog/ab-testing-voice-agents): A/B testing voice agents compares two versions on the same calls to find the winner. Learn to test prompts, models, and voices with real metrics. - [Voice agent prompt comparison: how to test prompt changes side by side](https://www.evalgent.com/blog/voice-agent-prompt-comparison): Voice agent prompt comparison runs two prompts on the same calls to find the winner. Learn side-by-side testing, catching regressions, and the tooling. - [The voice AI deployment gap: why demos win and production breaks](https://www.evalgent.com/blog/voice-ai-deployment-gap): The voice AI deployment gap is why agents that ace demos break in production. Learn what causes the pilot-to-production drop and how to close it. - [How to monitor AI voice agents in production](https://www.evalgent.com/blog/monitor-ai-voice-agents-production): How to monitor AI voice agents in production: the metrics to track, alerts to set, and how to detect drift across STT, LLM, and TTS before callers do. - [AI voice agent cost: what a voice agent really costs per minute](https://www.evalgent.com/blog/ai-voice-agent-cost): AI voice agent cost breaks into STT, LLM, TTS, telephony, and platform fees. Learn the 2026 per-minute price, what drives it, and how to budget and cut it. - [xAI Voice Agent: a guide to Grok Voice Agent Builder](https://www.evalgent.com/blog/xai-grok-voice-agent): xAI Voice Agent Builder turns a plain-language brief into a live Grok Voice agent in minutes. Learn its architecture, pricing, and features. - [xAI Voice Agent vs ElevenLabs: which voice platform should you choose?](https://www.evalgent.com/blog/xai-voice-agent-vs-elevenlabs): xAI Voice Agent vs ElevenLabs: single speech-to-speech model vs cascading pipeline, pricing, voice quality, LLM choice, and latency compared. - [Cascading vs speech-to-speech voice agents: which architecture should you choose?](https://www.evalgent.com/blog/cascading-vs-speech-to-speech-voice-agents): Cascading vs speech-to-speech voice agents: compare the pipeline and single-model architectures on latency, control, cost, and voice quality. - [How to test speech-to-speech voice agents](https://www.evalgent.com/blog/testing-speech-to-speech-voice-agents): Testing speech-to-speech voice agents is hard: no transcript to inspect. Learn why outcome-based, full-call testing is the only reliable approach.