xAI Voice Agent vs ElevenLabs: Grok Voice Agent Builder pricing per minute, features, and which to choose (2026)

On this page
Both products put a talking agent on a phone line. They get there in very different ways. This guide goes layer by layer: architecture, pricing per minute, voices, languages, turn-taking, tools, telephony, SDKs, observability, and compliance. Every product fact links to the vendor's own docs as of September 2026. Prices change often, so check current pricing before you commit.
Evalgent is an independent evaluation platform. We do not resell either vendor. The goal here is a neutral map, then a test plan you can run on your own calls.
xAI Grok Voice: xAI's realtime voice stack. It includes a speech-to-speech API (`grok-voice-think-fast-2.0`), a no-code Grok Voice Agent Builder in the xAI console, plus separate text-to-speech, speech-to-text, and voice cloning APIs.
ElevenLabs Agents (ElevenAgents): ElevenLabs' conversational agent platform, formerly branded Conversational AI. It chains speech-to-text, an LLM you choose, text-to-speech, and a proprietary turn-taking model.
Grok Voice vs ElevenLabs at a glance
Each row below is expanded later in the article.

| Dimension | xAI Grok Voice | ElevenLabs Agents |
|---|---|---|
| What you buy | Realtime API plus a no-code Agent Builder (beta) | Hosted agent platform with dashboard, API, and CLI |
| Architecture | Single speech-to-speech model | STT + your LLM + TTS + turn-taking model |
| Current model | `grok-voice-think-fast-2.0` (`grok-voice-latest`) | Your choice, such as GPT, Claude, Gemini, Qwen, or a custom LLM |
| List price | $0.08/min audio, $0.004 per text input | $0.08/min hosting, plus LLM and telephony |
| Built-in voices | 28 named voices in the API docs | 5,000+ voices in the library |
| Voice cloning | From a clip up to 120 seconds | Instant and Professional Voice Clones |
| Languages | 20 listed codes, "20+" claimed, auto-detect | 31 on Flash models; 70+ on Eleven v3 Conversational |
| Turn-taking controls | Server VAD threshold, silence, padding, idle timeout | Eagerness modes, turn timeout, soft timeout, interruptions |
| Tools | Functions, remote MCP, web search, X search, collections | Server tools, client tools, MCP, system tools |
| Telephony | Direct SIP (Twilio, Telnyx, Plivo, BYO); builder number | Native Twilio, SIP trunk, batch outbound calls |
| Session and concurrency | 120-min sessions; 10 concurrent sessions per team by default | 4 to 40 concurrent calls by plan; burst at $0.16/min |
| Observability | Call playback in builder; event stream via API | History, analysis, tests, experiments, OpenTelemetry |
| Documented compliance | SOC 2 Type II, HIPAA eligible, GDPR | HIPAA BAA (Enterprise, Zero Retention Mode), GDPR, EU residency |
Sources for each row appear in the sections below. The big split is simple. xAI sells one opinionated model. ElevenLabs sells a configurable platform.
What each product actually is
xAI: an API first, with a builder on top
xAI's voice offer has two front doors. The first is the Speech to Speech API, a WebSocket endpoint at `wss://api.x.ai/v1/realtime`. Audio goes in and audio comes out of one Grok model. The second is the Grok Voice Agent Builder, a no-code console that xAI still labels beta. xAI pitches it as a way to launch an agent in under two minutes.
The API is compatible with the OpenAI Realtime API. xAI documents a short migration: change the base URL, swap the key, pick a Grok model. A few events differ, and xAI lists them. That matters for lock-in, as we cover later.
ElevenLabs: a platform around your LLM
ElevenLabs Agents coordinates four components. A fine-tuned ASR model hears the caller. Your chosen LLM decides what to say. A low-latency TTS model speaks. A proprietary turn-taking model decides when. You configure it through a dashboard, a visual workflow builder, the API, or a CLI that treats agents as code.
This is the cascading design. It adds hand-offs, but each stage is swappable and inspectable. For the deeper trade-off, read our cascading vs speech-to-speech explainer.
Pricing per minute: Grok Voice Agent Builder vs ElevenLabs
Most searchers want this section. Here are the published numbers as of September 2026. Check current pricing before you budget.
xAI. The xAI pricing page lists Speech to Speech on `grok-voice-think-fast-2.0` at $0.08 per minute, or $4.80 per hour. Each text input message costs $0.004. The model page adds useful detail. Sessions using server VAD are billed for session duration, not just speech. Tool results sent back as `function_call_output` are not billed as text input. Server-side tools bill separately: collections search is $2.50 per 1,000 calls, and web search is $5 per 1,000 calls.
This corrects an older figure. Earlier write-ups, including ours, quoted about $0.05 per minute. That figure predates the current Think Fast 2.0 model. Today's documented API rate is $0.08.
The Agent Builder includes a free phone number and 30 free voice clones. Per xAI's Voice Agent Builder announcement, agents bill at the API rate with no separate platform fee, and telephony on a free provisioned number adds $0.01 per minute (as of September 2026; check current pricing).
ElevenLabs. The ElevenAgents pricing page lists $0.08 per call minute on every plan. Burst minutes above your concurrency limit cost $0.16. Text messages cost $0.003 each. The LLM is billed by usage and varies by model. Telephony is billed at cost. Plans bundle minutes: 15 on Free, 1,238 on Pro ($99/month), and 12,375 on Business ($990/month).
An illustrative cost per 1,000 minutes
List prices hide the real bill. The model below adds the other meters. Every number marked illustrative is our assumption, not a vendor quote.

| Line item (per 1,000 minutes) | xAI Grok Voice API | ElevenLabs Agents |
|---|---|---|
| Platform or model minutes | $80.00 (published) | $80.00 (published) |
| LLM tokens | $0 (included in the model) | $15.00 (illustrative, mid-size model) |
| Telephony, SIP or Twilio | $10.00 (illustrative) | $10.00 (illustrative) |
| Knowledge lookups | $1.67 (about 667 collection searches) | $0 (RAG included in hosting) |
| Illustrative total | about $92 | about $105 |
| Illustrative per minute | about $0.092 | about $0.105 |
Three things swing this model more than the headline rate. First, LLM choice: a large reasoning model on ElevenLabs can cost more than the platform minute. Second, idle time: xAI bills server-VAD sessions for their full duration, so hold music and long silences cost money. Third, concurrency: ElevenLabs burst minutes double the rate during spikes. Our voice agent pricing hidden costs guide covers more traps, and AI voice agent cost gives the wider market view.
Architecture: single speech-to-speech vs cascading pipeline
xAI runs one model from audio to audio. Nothing is stitched together on each turn. That design tends to keep latency low and prosody natural. It also means you cannot swap the reasoning model. You can tune it, though. The `reasoning.effort` session setting accepts `"high"` or `"none"`, and xAI defaults to high.
ElevenLabs runs a pipeline. Every turn passes through ASR, the LLM, and TTS. You can see and change each part. You can also add a backup LLM, which ElevenLabs strongly recommends for production. The cost is more hand-offs and settings. Our best LLM for voice agents guide helps with model picks.
Neither design is "better" in the abstract. Speech-to-speech models are harder to debug, because there is no clean text hand-off to inspect. Pipelines are easier to debug but can accumulate latency. Our guide to testing speech-to-speech voice agents explains why the test approach differs.
Voices, cloning, and voice quality
xAI. The voice overview lists 28 named built-in voices. Examples include `eve` (the default), `ara`, `leo`, `rex`, and `celeste`. xAI says every voice can speak every supported language. Custom voices come from a reference clip of up to 120 seconds. The resulting `voice_id` works in the realtime API. A `replace` setting fixes mispronounced brand names before audio is generated. An `audio.output.speed` setting ranges from 0.7 to 1.5.
ElevenLabs. The agent platform offers 5,000+ voices. The new Expressive mode runs on Eleven v3 Conversational. It adapts tone to the caller and accepts tags like `[laughs]` or `[sighs]`. ElevenLabs says it costs the same as other agent TTS models. One caveat is documented: v3 Conversational does not preserve Professional Voice Clone characteristics well. If your cloned brand voice matters, ElevenLabs suggests Flash v2 instead.
So library size clearly favors ElevenLabs. Perceived quality is a different question. It depends on your script, your callers, and the phone codec. Run a blind listening test on 8 kHz audio, as our TTS evaluation guide describes.
Languages and mid-call switching
xAI's speech-to-speech docs list 20 language codes and claim "20+" languages. The list includes English, Spanish (Mexico and Spain), Portuguese (Brazil and Portugal), French, German, Hindi, Japanese, Korean, and three Arabic variants. The model auto-detects the input language. You can bias recognition with a `language_hint` and up to 100 `keyterms`, both changeable mid-session. The Agent Builder page markets "25+ languages". We use the API docs figure here.
ElevenLabs Agents supports 31 languages when you select "All" on the default models. Eleven v3 Conversational expands TTS to 70+ languages. Automatic switching uses a language detection system tool. The language page also notes that preset language selection is fixed for the call. Test switching behavior directly if bilingual callers matter.
Latency, turn-taking, and interruptions
Both vendors claim fast responses. xAI states sub-second latency. Neither claim tells you how your agent performs with your prompt, tools, and phone carrier. Measure it.
The controls differ in useful ways:
| Control | xAI Grok Voice | ElevenLabs Agents |
|---|---|---|
| End-of-turn | `server_vad` with threshold 0.1 to 0.9 (default 0.85) | Turn eagerness: eager, normal, or patient |
| Pause tolerance | `silence_duration_ms`, 0 to 10,000 ms | Turn timeout, 1 to 30 seconds |
| Clipped word starts | `prefix_padding_ms` (default 333) | Handled by the turn-taking model |
| Filler while thinking | Not a documented setting | Soft timeout, 0.5 to 8 seconds |
| Barge-in | On with server VAD; per-message `interruptible: false` | Toggle interruptions on or off |
| Re-engage silent caller | `idle_timeout_ms` | Turn timeout prompts the caller |
Sources: xAI session parameters and ElevenLabs conversation flow.
A simplified xAI session setup looks like this:
# Simplified: configure a Grok voice session for phone audio
session = {
"type": "session.update",
"session": {
"voice": "ara",
"instructions": "You are a scheduling agent for Sunrise Dental.",
"reasoning": {"effort": "none"}, # trade depth for speed
"turn_detection": {
"type": "server_vad",
"threshold": 0.85,
"silence_duration_ms": 700,
"prefix_padding_ms": 333,
},
"audio": {
"input": {"format": {"type": "audio/pcmu"}},
"output": {"format": {"type": "audio/pcmu"}},
},
"tools": [{"type": "file_search", "vector_store_ids": ["kb_123"]}],
},
}The ElevenLabs equivalent lives in agent configuration JSON. You can pull and push it with the CLI:
{
"conversation_config": {
"turn": {
"turn_eagerness": "patient",
"turn_timeout": 7,
"soft_timeout_config": { "timeout_seconds": 3.0, "message": "One moment." }
},
"tts": { "model_id": "eleven_v3_conversational" },
"conversation": { "max_duration_seconds": 1200 }
}
}The failure modes differ too. A high VAD threshold on xAI can miss quiet callers. A short silence window cuts off people reading card numbers. On ElevenLabs, "eager" mode can talk over slow speakers, and a slow LLM shows up as dead air. Our guides on barge-in and latency cover how to measure both.
Tools, function calling, and knowledge bases
xAI. The realtime session accepts five tool types: `function`, `file_search`, `web_search`, `x_search`, and `mcp`. Server-side tools run inside xAI, so you do not handle their results. Custom functions return through a `function_call_output` item. xAI supports parallel tool calls. For grounding, `file_search` queries your uploaded Collections. The Agent Builder adds ready connectors, including Gmail, Google Calendar, Outlook, Linear, and Notion.
ElevenLabs. Agents support server tools (webhooks), client tools, MCP tools, and system tools such as language detection and skip turn. The knowledge base accepts documents and URLs, with RAG you can enable per agent. A visual workflow builder handles multi-step flows with different settings per node.
In both cases, the risk is the same. Tool calls fail silently, return stale data, or fire twice. Read our tool calling guide before trusting either in production.
Telephony: SIP, Twilio, and phone numbers
xAI. Phone calls use Direct SIP. You register your own number with `origin: "byo_trunk"`. xAI does not provision numbers through the API. Your carrier routes calls to `sip.voice.x.ai`, and xAI sends a signed webhook. Your server then joins the call over WebSocket. The docs give setup steps for Twilio Elastic SIP Trunking, Telnyx, and Plivo. Call control includes `refer` transfers, `hangup`, and DTMF keypress capture. The Agent Builder separately offers a free number for quick starts.
# Simplified: register a BYO-trunk number with xAI (from xAI SIP docs)
curl -X POST "https://api.x.ai/v2/phone-numbers" \
-H "Authorization: Bearer $XAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"origin": "byo_trunk", "phone_number": "+18005550199",
"webhook": {"url": "https://example.com/xai/sip-webhook"},
"sip_auth": {"allowed_addresses": ["203.0.113.0/24"]}}'ElevenLabs. Agents offer a native Twilio integration, generic SIP trunking, and batch outbound calling. Telephony is billed at cost. That makes outbound campaigns easier to launch without custom code.
For the wider carrier picture, see best telephony for voice agents. If you want to keep options open, own the number on your own trunk. Then either platform can sit behind it.
SDKs, deployment, and observability
xAI is closer to raw infrastructure. You get the WebSocket API, ephemeral tokens for browsers, and cookbook demos for web, WebRTC, Twilio, and iOS. LiveKit and Pipecat both document Grok integrations. Because the API mirrors OpenAI Realtime, many existing clients work with a URL change. The builder adds in-browser testing and call playback. Beyond that, logging and analytics are yours to build.
ElevenLabs ships more around the agent. There are Python and JavaScript SDKs, React, React Native, Swift, and Kotlin SDKs, and an embeddable widget. Operations features include conversation history, success evaluation, data collection, automated tests, live A/B experiments, post-call webhooks, and OpenTelemetry traces.
Built-in analytics are useful, but they grade the vendor's own work. Treat them as a starting point, not an independent audit. Our ElevenLabs voice agent testing guide shows how to extend them.
Compliance and data handling
Only documented claims appear here. Always confirm with each vendor's legal team.
xAI. The voice overview lists SOC 2 Type II, HIPAA eligibility with a BAA, GDPR compliance, data residency options, and SAML SSO. It also states that voice audio is processed in real time and never stored or used for training. The model page lists the realtime cluster as us-east-1.
ElevenLabs. The HIPAA page says BAAs are Enterprise-only and require Zero Retention Mode. In that mode, transcripts, audio, and tool results are not stored. Only allowlisted LLMs are available unless you bring your own under a separate BAA. ElevenLabs also offers EU data residency for Enterprise customers and publishes a Trust Center.
The practical gap: Zero Retention Mode limits ElevenLabs' built-in analytics. Healthcare teams will need their own evaluation records. Our compliance officer's guide lists the questions to ask.
Which to choose, by use case
| Use case | Leans toward | Why |
|---|---|---|
| High-volume inbound support, English-first | xAI | One bundled meter; strong tool calling; simple stack |
| Brand-led experiences where the voice is the product | ElevenLabs | 5,000+ voices; expressive v3 delivery |
| Teams with a required LLM (policy, cost, or quality) | ElevenLabs | Native model choice plus custom LLM endpoints |
| Existing OpenAI Realtime code | xAI | Compatible API; minimal migration |
| Outbound campaigns at scale | ElevenLabs | Batch calling and native Twilio |
| Healthcare with PHI | Either, with care | Both document HIPAA paths; check BAA scope and retention |
| Many languages beyond the top 20 | ElevenLabs | 70+ on v3 Conversational |
| Non-technical team, fast pilot | xAI builder or ElevenLabs dashboard | Both offer no-code paths; the xAI builder is beta |
These are leanings, not verdicts. Your own test data decides.
Migration and lock-in
Lock-in lives in four places: prompts, tools, numbers, and voices.
- Prompts move easily, but speech-to-speech and cascaded models respond differently to the same prompt. Budget time to retune.
- Tools port well if you expose them over MCP or plain HTTP. Both platforms accept MCP.
- Phone numbers stay portable if they live on your own SIP trunk.
- Voice clones do not transfer. You must re-clone on the new platform, and the result will sound different.
xAI's OpenAI Realtime compatibility lowers switching costs between realtime APIs. ElevenLabs' custom LLM option accepts any OpenAI-compatible endpoint. That lets you change the brain without leaving the platform. See voice AI vendor lock-in for a fuller checklist.
How to run a fair side-by-side test on your own calls
A fair test holds everything constant except the platform. Our voice agent bake-off guide covers the procurement side.
1. Freeze one spec. Write one agent brief: same greeting, same policies, same tools, same knowledge files. Port it to both platforms.
2. Match the audio path. Route both agents through the same carrier and codec, ideally G.711 over SIP. Browser tests flatter everyone.
3. Pin versions. Use `grok-voice-think-fast-2.0`, not `grok-voice-latest`, and pin the ElevenLabs LLM and TTS model.
4. Build a scenario set. Aim for 50 to 100 scenarios from real call logs. Include accents, noise, interruptions, long pauses, and spelled-out IDs.
5. Run synthetic callers at volume. Run each scenario several times per platform. Voice agents are not deterministic, so single runs mislead.
6. Measure the same metrics. Track task success, tool-call accuracy, time to first audio, interruption recovery, and cost per resolved call.
7. Test failure paths. Force a tool timeout and a failed transfer. Watch what each agent says.
8. Compare cost on real minutes. Multiply measured minutes by current list prices, including idle time and LLM tokens.
9. Blind-review a sample. Have reviewers score recordings without knowing which platform made them.
A simple scoring pass might look like this:
# Illustrative: compare two platforms on identical scenario runs
import statistics as st
def summarize(runs):
return {
"task_success": sum(r["success"] for r in runs) / len(runs),
"p50_first_audio_ms": st.median(r["first_audio_ms"] for r in runs),
"p95_first_audio_ms": sorted(r["first_audio_ms"] for r in runs)[int(0.95 * len(runs)) - 1],
"cost_per_success": sum(r["cost_usd"] for r in runs)
/ max(1, sum(r["success"] for r in runs)),
}
report = {"xai": summarize(xai_runs), "elevenlabs": summarize(eleven_runs)}Cost per successful call is the number that settles most debates. A cheaper minute that fails more often costs more. Our guide to comparing voice agents on the same test cases goes deeper on scenario design.
Where Evalgent fits
Neither platform proves its agent works with your callers. Both ship some testing, but a vendor grading itself is not independent evidence. Evalgent runs the same synthetic callers against agents on Grok Voice, ElevenLabs, or any other stack. Scenarios reproduce noisy, accented, and off-script calls. Metrics score task completion, latency, and policy adherence per cohort. Reviews let your team hear exactly where each agent struggled. The AI voice agent testing pillar explains the method, and the xAI Voice Agent guide covers Grok in more depth.
Frequently asked questions
How much does the Grok Voice Agent Builder cost per minute?
As of September 2026, xAI's speech-to-speech API lists $0.08 per minute on `grok-voice-think-fast-2.0`, plus $0.004 per text input message. Server-side tools such as collections search bill separately. Builder agents bill at that same API rate with no platform fee, and a free provisioned number adds $0.01 per minute of telephony.
What is Grok Voice Think Fast 2.0 pricing?
Grok Voice Think Fast 2.0 is xAI's current flagship voice model, also reachable as `grok-voice-latest`. xAI prices it at $0.08 per audio minute, or $4.80 per hour, as of September 2026. Sessions using server VAD bill for full session duration. Earlier $0.05 figures applied to the previous Think Fast 1.0 model.
Is Grok voice cheaper than ElevenLabs?
Often, but not by much. Both list $0.08 per minute as of September 2026. xAI includes the model in that rate. ElevenLabs adds LLM tokens and telephony at cost. In our illustrative model, xAI lands near $0.09 per minute and ElevenLabs near $0.10. Idle time, LLM choice, and burst pricing can flip the result.
How do I set up a Grok voice agent to answer phone calls?
Use the Agent Builder's free number for a quick test. For production, register your own number with xAI's phone-numbers API as a BYO trunk. Then point your Twilio, Telnyx, or Plivo SIP trunk at `sip.voice.x.ai`. xAI sends a signed webhook, and your server joins the call over WebSocket.
Which has better voices, Grok or ElevenLabs?
ElevenLabs offers far more choice, with 5,000+ voices and expressive Eleven v3 Conversational delivery. xAI documents 28 built-in voices plus cloning from short clips. "Better" depends on your callers and phone audio. Run a blind test at 8 kHz before deciding, because studio demos rarely match real phone quality.
Which supports more languages, Grok or ElevenLabs?
ElevenLabs covers more. Its agents support 31 languages on default models and 70+ with Eleven v3 Conversational. xAI's speech-to-speech docs list 20 language codes and claim 20+ languages, with automatic detection. For bilingual callers, test mid-call switching on both, since behavior varies by configuration.
Can I use my own LLM with Grok voice or ElevenLabs?
Not with xAI's speech-to-speech model. The Grok model is fixed, though you can set reasoning effort. ElevenLabs natively supports GPT, Claude, Gemini, and hosted Qwen models. It also accepts any OpenAI-compatible custom LLM endpoint. If model choice is a hard requirement, ElevenLabs is the only one of the two that offers it.
Do you still need to test a Grok or ElevenLabs voice agent?
Yes. Both platforms help you build and launch, but neither proves your agent handles your callers. Accents, noise, interruptions, and tool failures break agents on every stack. Independent testing with synthetic callers, such as Evalgent, measures task success and latency before real customers find the gaps.
The bottom line
xAI Grok Voice and ElevenLabs Agents both list $0.08 per minute in September 2026, but xAI bundles the model while ElevenLabs bills the LLM and telephony separately. The right choice is the one that wins on cost per successful call in a fair test on your own traffic.
Book a demo to run that side-by-side test with Evalgent: Book a demo.
Related Articles

Why AI voice agents fail in production: 9 failure layers, how to detect each, and how to prevent them
Voice agents fail in production across 9 layers, from 8 kHz audio to silent tool errors. Symptoms, root causes, alerts, and fixes in one master table.
Read more
Voice agent regression testing: how to detect regressions from model updates you don't control
Voice agent regression testing for changes you don't control: pin each layer, run a nightly pinned-vs-latest canary, and use paired stats to detect drops.
Read more