Evalgent
Back to Blog
Voice AI Evaluation

xAI Voice Agent vs ElevenLabs: Grok Voice Agent Builder pricing per minute, features, and which to choose (2026)

Deepesh Jayal
17 min read
xAI Voice Agent vs ElevenLabs: Grok Voice Agent Builder pricing per minute, features, and which to choose (2026)
On this page

Both products put a talking agent on a phone line. They get there in very different ways. This guide goes layer by layer: architecture, pricing per minute, voices, languages, turn-taking, tools, telephony, SDKs, observability, and compliance. Every product fact links to the vendor's own docs as of September 2026. Prices change often, so check current pricing before you commit.

Evalgent is an independent evaluation platform. We do not resell either vendor. The goal here is a neutral map, then a test plan you can run on your own calls.

xAI Grok Voice: xAI's realtime voice stack. It includes a speech-to-speech API (`grok-voice-think-fast-2.0`), a no-code Grok Voice Agent Builder in the xAI console, plus separate text-to-speech, speech-to-text, and voice cloning APIs.

ElevenLabs Agents (ElevenAgents): ElevenLabs' conversational agent platform, formerly branded Conversational AI. It chains speech-to-text, an LLM you choose, text-to-speech, and a proprietary turn-taking model.

Grok Voice vs ElevenLabs at a glance

Each row below is expanded later in the article.

xAI Grok voice agents versus ElevenLabs Agents compared side by side: model approach, voices, languages, turn-taking, tools, telephony and pricing model
DimensionxAI Grok VoiceElevenLabs Agents
What you buyRealtime API plus a no-code Agent Builder (beta)Hosted agent platform with dashboard, API, and CLI
ArchitectureSingle speech-to-speech modelSTT + your LLM + TTS + turn-taking model
Current model`grok-voice-think-fast-2.0` (`grok-voice-latest`)Your choice, such as GPT, Claude, Gemini, Qwen, or a custom LLM
List price$0.08/min audio, $0.004 per text input$0.08/min hosting, plus LLM and telephony
Built-in voices28 named voices in the API docs5,000+ voices in the library
Voice cloningFrom a clip up to 120 secondsInstant and Professional Voice Clones
Languages20 listed codes, "20+" claimed, auto-detect31 on Flash models; 70+ on Eleven v3 Conversational
Turn-taking controlsServer VAD threshold, silence, padding, idle timeoutEagerness modes, turn timeout, soft timeout, interruptions
ToolsFunctions, remote MCP, web search, X search, collectionsServer tools, client tools, MCP, system tools
TelephonyDirect SIP (Twilio, Telnyx, Plivo, BYO); builder numberNative Twilio, SIP trunk, batch outbound calls
Session and concurrency120-min sessions; 10 concurrent sessions per team by default4 to 40 concurrent calls by plan; burst at $0.16/min
ObservabilityCall playback in builder; event stream via APIHistory, analysis, tests, experiments, OpenTelemetry
Documented complianceSOC 2 Type II, HIPAA eligible, GDPRHIPAA BAA (Enterprise, Zero Retention Mode), GDPR, EU residency

Sources for each row appear in the sections below. The big split is simple. xAI sells one opinionated model. ElevenLabs sells a configurable platform.

What each product actually is

xAI: an API first, with a builder on top

xAI's voice offer has two front doors. The first is the Speech to Speech API, a WebSocket endpoint at `wss://api.x.ai/v1/realtime`. Audio goes in and audio comes out of one Grok model. The second is the Grok Voice Agent Builder, a no-code console that xAI still labels beta. xAI pitches it as a way to launch an agent in under two minutes.

The API is compatible with the OpenAI Realtime API. xAI documents a short migration: change the base URL, swap the key, pick a Grok model. A few events differ, and xAI lists them. That matters for lock-in, as we cover later.

ElevenLabs: a platform around your LLM

ElevenLabs Agents coordinates four components. A fine-tuned ASR model hears the caller. Your chosen LLM decides what to say. A low-latency TTS model speaks. A proprietary turn-taking model decides when. You configure it through a dashboard, a visual workflow builder, the API, or a CLI that treats agents as code.

This is the cascading design. It adds hand-offs, but each stage is swappable and inspectable. For the deeper trade-off, read our cascading vs speech-to-speech explainer.

Pricing per minute: Grok Voice Agent Builder vs ElevenLabs

Most searchers want this section. Here are the published numbers as of September 2026. Check current pricing before you budget.

xAI. The xAI pricing page lists Speech to Speech on `grok-voice-think-fast-2.0` at $0.08 per minute, or $4.80 per hour. Each text input message costs $0.004. The model page adds useful detail. Sessions using server VAD are billed for session duration, not just speech. Tool results sent back as `function_call_output` are not billed as text input. Server-side tools bill separately: collections search is $2.50 per 1,000 calls, and web search is $5 per 1,000 calls.

This corrects an older figure. Earlier write-ups, including ours, quoted about $0.05 per minute. That figure predates the current Think Fast 2.0 model. Today's documented API rate is $0.08.

The Agent Builder includes a free phone number and 30 free voice clones. Per xAI's Voice Agent Builder announcement, agents bill at the API rate with no separate platform fee, and telephony on a free provisioned number adds $0.01 per minute (as of September 2026; check current pricing).

ElevenLabs. The ElevenAgents pricing page lists $0.08 per call minute on every plan. Burst minutes above your concurrency limit cost $0.16. Text messages cost $0.003 each. The LLM is billed by usage and varies by model. Telephony is billed at cost. Plans bundle minutes: 15 on Free, 1,238 on Pro ($99/month), and 12,375 on Business ($990/month).

$0.08
xAI speech-to-speech per minute, model included
$0.08
ElevenLabs Agents per minute, before LLM and telephony
$0.16
ElevenLabs burst minute above your concurrency limit
120 min
Max xAI realtime session length

An illustrative cost per 1,000 minutes

List prices hide the real bill. The model below adds the other meters. Every number marked illustrative is our assumption, not a vendor quote.

An illustrative cost-per-minute model for xAI and ElevenLabs voice agents: platform or model minutes, telephony, and LLM or tool costs added up per 1,000 minutes
Line item (per 1,000 minutes)xAI Grok Voice APIElevenLabs Agents
Platform or model minutes$80.00 (published)$80.00 (published)
LLM tokens$0 (included in the model)$15.00 (illustrative, mid-size model)
Telephony, SIP or Twilio$10.00 (illustrative)$10.00 (illustrative)
Knowledge lookups$1.67 (about 667 collection searches)$0 (RAG included in hosting)
Illustrative totalabout $92about $105
Illustrative per minuteabout $0.092about $0.105

Three things swing this model more than the headline rate. First, LLM choice: a large reasoning model on ElevenLabs can cost more than the platform minute. Second, idle time: xAI bills server-VAD sessions for their full duration, so hold music and long silences cost money. Third, concurrency: ElevenLabs burst minutes double the rate during spikes. Our voice agent pricing hidden costs guide covers more traps, and AI voice agent cost gives the wider market view.

Architecture: single speech-to-speech vs cascading pipeline

xAI runs one model from audio to audio. Nothing is stitched together on each turn. That design tends to keep latency low and prosody natural. It also means you cannot swap the reasoning model. You can tune it, though. The `reasoning.effort` session setting accepts `"high"` or `"none"`, and xAI defaults to high.

ElevenLabs runs a pipeline. Every turn passes through ASR, the LLM, and TTS. You can see and change each part. You can also add a backup LLM, which ElevenLabs strongly recommends for production. The cost is more hand-offs and settings. Our best LLM for voice agents guide helps with model picks.

Neither design is "better" in the abstract. Speech-to-speech models are harder to debug, because there is no clean text hand-off to inspect. Pipelines are easier to debug but can accumulate latency. Our guide to testing speech-to-speech voice agents explains why the test approach differs.

Voices, cloning, and voice quality

xAI. The voice overview lists 28 named built-in voices. Examples include `eve` (the default), `ara`, `leo`, `rex`, and `celeste`. xAI says every voice can speak every supported language. Custom voices come from a reference clip of up to 120 seconds. The resulting `voice_id` works in the realtime API. A `replace` setting fixes mispronounced brand names before audio is generated. An `audio.output.speed` setting ranges from 0.7 to 1.5.

ElevenLabs. The agent platform offers 5,000+ voices. The new Expressive mode runs on Eleven v3 Conversational. It adapts tone to the caller and accepts tags like `[laughs]` or `[sighs]`. ElevenLabs says it costs the same as other agent TTS models. One caveat is documented: v3 Conversational does not preserve Professional Voice Clone characteristics well. If your cloned brand voice matters, ElevenLabs suggests Flash v2 instead.

So library size clearly favors ElevenLabs. Perceived quality is a different question. It depends on your script, your callers, and the phone codec. Run a blind listening test on 8 kHz audio, as our TTS evaluation guide describes.

Languages and mid-call switching

xAI's speech-to-speech docs list 20 language codes and claim "20+" languages. The list includes English, Spanish (Mexico and Spain), Portuguese (Brazil and Portugal), French, German, Hindi, Japanese, Korean, and three Arabic variants. The model auto-detects the input language. You can bias recognition with a `language_hint` and up to 100 `keyterms`, both changeable mid-session. The Agent Builder page markets "25+ languages". We use the API docs figure here.

ElevenLabs Agents supports 31 languages when you select "All" on the default models. Eleven v3 Conversational expands TTS to 70+ languages. Automatic switching uses a language detection system tool. The language page also notes that preset language selection is fixed for the call. Test switching behavior directly if bilingual callers matter.

Latency, turn-taking, and interruptions

Both vendors claim fast responses. xAI states sub-second latency. Neither claim tells you how your agent performs with your prompt, tools, and phone carrier. Measure it.

The controls differ in useful ways:

ControlxAI Grok VoiceElevenLabs Agents
End-of-turn`server_vad` with threshold 0.1 to 0.9 (default 0.85)Turn eagerness: eager, normal, or patient
Pause tolerance`silence_duration_ms`, 0 to 10,000 msTurn timeout, 1 to 30 seconds
Clipped word starts`prefix_padding_ms` (default 333)Handled by the turn-taking model
Filler while thinkingNot a documented settingSoft timeout, 0.5 to 8 seconds
Barge-inOn with server VAD; per-message `interruptible: false`Toggle interruptions on or off
Re-engage silent caller`idle_timeout_ms`Turn timeout prompts the caller

Sources: xAI session parameters and ElevenLabs conversation flow.

A simplified xAI session setup looks like this:

# Simplified: configure a Grok voice session for phone audio
session = {
    "type": "session.update",
    "session": {
        "voice": "ara",
        "instructions": "You are a scheduling agent for Sunrise Dental.",
        "reasoning": {"effort": "none"},   # trade depth for speed
        "turn_detection": {
            "type": "server_vad",
            "threshold": 0.85,
            "silence_duration_ms": 700,
            "prefix_padding_ms": 333,
        },
        "audio": {
            "input": {"format": {"type": "audio/pcmu"}},
            "output": {"format": {"type": "audio/pcmu"}},
        },
        "tools": [{"type": "file_search", "vector_store_ids": ["kb_123"]}],
    },
}

The ElevenLabs equivalent lives in agent configuration JSON. You can pull and push it with the CLI:

{
  "conversation_config": {
    "turn": {
      "turn_eagerness": "patient",
      "turn_timeout": 7,
      "soft_timeout_config": { "timeout_seconds": 3.0, "message": "One moment." }
    },
    "tts": { "model_id": "eleven_v3_conversational" },
    "conversation": { "max_duration_seconds": 1200 }
  }
}

The failure modes differ too. A high VAD threshold on xAI can miss quiet callers. A short silence window cuts off people reading card numbers. On ElevenLabs, "eager" mode can talk over slow speakers, and a slow LLM shows up as dead air. Our guides on barge-in and latency cover how to measure both.

Tools, function calling, and knowledge bases

xAI. The realtime session accepts five tool types: `function`, `file_search`, `web_search`, `x_search`, and `mcp`. Server-side tools run inside xAI, so you do not handle their results. Custom functions return through a `function_call_output` item. xAI supports parallel tool calls. For grounding, `file_search` queries your uploaded Collections. The Agent Builder adds ready connectors, including Gmail, Google Calendar, Outlook, Linear, and Notion.

ElevenLabs. Agents support server tools (webhooks), client tools, MCP tools, and system tools such as language detection and skip turn. The knowledge base accepts documents and URLs, with RAG you can enable per agent. A visual workflow builder handles multi-step flows with different settings per node.

In both cases, the risk is the same. Tool calls fail silently, return stale data, or fire twice. Read our tool calling guide before trusting either in production.

Telephony: SIP, Twilio, and phone numbers

xAI. Phone calls use Direct SIP. You register your own number with `origin: "byo_trunk"`. xAI does not provision numbers through the API. Your carrier routes calls to `sip.voice.x.ai`, and xAI sends a signed webhook. Your server then joins the call over WebSocket. The docs give setup steps for Twilio Elastic SIP Trunking, Telnyx, and Plivo. Call control includes `refer` transfers, `hangup`, and DTMF keypress capture. The Agent Builder separately offers a free number for quick starts.

# Simplified: register a BYO-trunk number with xAI (from xAI SIP docs)
curl -X POST "https://api.x.ai/v2/phone-numbers" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"origin": "byo_trunk", "phone_number": "+18005550199",
       "webhook": {"url": "https://example.com/xai/sip-webhook"},
       "sip_auth": {"allowed_addresses": ["203.0.113.0/24"]}}'

ElevenLabs. Agents offer a native Twilio integration, generic SIP trunking, and batch outbound calling. Telephony is billed at cost. That makes outbound campaigns easier to launch without custom code.

For the wider carrier picture, see best telephony for voice agents. If you want to keep options open, own the number on your own trunk. Then either platform can sit behind it.

SDKs, deployment, and observability

xAI is closer to raw infrastructure. You get the WebSocket API, ephemeral tokens for browsers, and cookbook demos for web, WebRTC, Twilio, and iOS. LiveKit and Pipecat both document Grok integrations. Because the API mirrors OpenAI Realtime, many existing clients work with a URL change. The builder adds in-browser testing and call playback. Beyond that, logging and analytics are yours to build.

ElevenLabs ships more around the agent. There are Python and JavaScript SDKs, React, React Native, Swift, and Kotlin SDKs, and an embeddable widget. Operations features include conversation history, success evaluation, data collection, automated tests, live A/B experiments, post-call webhooks, and OpenTelemetry traces.

Built-in analytics are useful, but they grade the vendor's own work. Treat them as a starting point, not an independent audit. Our ElevenLabs voice agent testing guide shows how to extend them.

Compliance and data handling

Only documented claims appear here. Always confirm with each vendor's legal team.

xAI. The voice overview lists SOC 2 Type II, HIPAA eligibility with a BAA, GDPR compliance, data residency options, and SAML SSO. It also states that voice audio is processed in real time and never stored or used for training. The model page lists the realtime cluster as us-east-1.

ElevenLabs. The HIPAA page says BAAs are Enterprise-only and require Zero Retention Mode. In that mode, transcripts, audio, and tool results are not stored. Only allowlisted LLMs are available unless you bring your own under a separate BAA. ElevenLabs also offers EU data residency for Enterprise customers and publishes a Trust Center.

The practical gap: Zero Retention Mode limits ElevenLabs' built-in analytics. Healthcare teams will need their own evaluation records. Our compliance officer's guide lists the questions to ask.

Which to choose, by use case

Use caseLeans towardWhy
High-volume inbound support, English-firstxAIOne bundled meter; strong tool calling; simple stack
Brand-led experiences where the voice is the productElevenLabs5,000+ voices; expressive v3 delivery
Teams with a required LLM (policy, cost, or quality)ElevenLabsNative model choice plus custom LLM endpoints
Existing OpenAI Realtime codexAICompatible API; minimal migration
Outbound campaigns at scaleElevenLabsBatch calling and native Twilio
Healthcare with PHIEither, with careBoth document HIPAA paths; check BAA scope and retention
Many languages beyond the top 20ElevenLabs70+ on v3 Conversational
Non-technical team, fast pilotxAI builder or ElevenLabs dashboardBoth offer no-code paths; the xAI builder is beta

These are leanings, not verdicts. Your own test data decides.

Migration and lock-in

Lock-in lives in four places: prompts, tools, numbers, and voices.

  • Prompts move easily, but speech-to-speech and cascaded models respond differently to the same prompt. Budget time to retune.
  • Tools port well if you expose them over MCP or plain HTTP. Both platforms accept MCP.
  • Phone numbers stay portable if they live on your own SIP trunk.
  • Voice clones do not transfer. You must re-clone on the new platform, and the result will sound different.

xAI's OpenAI Realtime compatibility lowers switching costs between realtime APIs. ElevenLabs' custom LLM option accepts any OpenAI-compatible endpoint. That lets you change the brain without leaving the platform. See voice AI vendor lock-in for a fuller checklist.

How to run a fair side-by-side test on your own calls

A fair test holds everything constant except the platform. Our voice agent bake-off guide covers the procurement side.

1. Freeze one spec. Write one agent brief: same greeting, same policies, same tools, same knowledge files. Port it to both platforms.

2. Match the audio path. Route both agents through the same carrier and codec, ideally G.711 over SIP. Browser tests flatter everyone.

3. Pin versions. Use `grok-voice-think-fast-2.0`, not `grok-voice-latest`, and pin the ElevenLabs LLM and TTS model.

4. Build a scenario set. Aim for 50 to 100 scenarios from real call logs. Include accents, noise, interruptions, long pauses, and spelled-out IDs.

5. Run synthetic callers at volume. Run each scenario several times per platform. Voice agents are not deterministic, so single runs mislead.

6. Measure the same metrics. Track task success, tool-call accuracy, time to first audio, interruption recovery, and cost per resolved call.

7. Test failure paths. Force a tool timeout and a failed transfer. Watch what each agent says.

8. Compare cost on real minutes. Multiply measured minutes by current list prices, including idle time and LLM tokens.

9. Blind-review a sample. Have reviewers score recordings without knowing which platform made them.

A simple scoring pass might look like this:

# Illustrative: compare two platforms on identical scenario runs
import statistics as st

def summarize(runs):
    return {
        "task_success": sum(r["success"] for r in runs) / len(runs),
        "p50_first_audio_ms": st.median(r["first_audio_ms"] for r in runs),
        "p95_first_audio_ms": sorted(r["first_audio_ms"] for r in runs)[int(0.95 * len(runs)) - 1],
        "cost_per_success": sum(r["cost_usd"] for r in runs)
                            / max(1, sum(r["success"] for r in runs)),
    }

report = {"xai": summarize(xai_runs), "elevenlabs": summarize(eleven_runs)}

Cost per successful call is the number that settles most debates. A cheaper minute that fails more often costs more. Our guide to comparing voice agents on the same test cases goes deeper on scenario design.

Where Evalgent fits

Neither platform proves its agent works with your callers. Both ship some testing, but a vendor grading itself is not independent evidence. Evalgent runs the same synthetic callers against agents on Grok Voice, ElevenLabs, or any other stack. Scenarios reproduce noisy, accented, and off-script calls. Metrics score task completion, latency, and policy adherence per cohort. Reviews let your team hear exactly where each agent struggled. The AI voice agent testing pillar explains the method, and the xAI Voice Agent guide covers Grok in more depth.

Frequently asked questions

How much does the Grok Voice Agent Builder cost per minute?

As of September 2026, xAI's speech-to-speech API lists $0.08 per minute on `grok-voice-think-fast-2.0`, plus $0.004 per text input message. Server-side tools such as collections search bill separately. Builder agents bill at that same API rate with no platform fee, and a free provisioned number adds $0.01 per minute of telephony.

What is Grok Voice Think Fast 2.0 pricing?

Grok Voice Think Fast 2.0 is xAI's current flagship voice model, also reachable as `grok-voice-latest`. xAI prices it at $0.08 per audio minute, or $4.80 per hour, as of September 2026. Sessions using server VAD bill for full session duration. Earlier $0.05 figures applied to the previous Think Fast 1.0 model.

Is Grok voice cheaper than ElevenLabs?

Often, but not by much. Both list $0.08 per minute as of September 2026. xAI includes the model in that rate. ElevenLabs adds LLM tokens and telephony at cost. In our illustrative model, xAI lands near $0.09 per minute and ElevenLabs near $0.10. Idle time, LLM choice, and burst pricing can flip the result.

How do I set up a Grok voice agent to answer phone calls?

Use the Agent Builder's free number for a quick test. For production, register your own number with xAI's phone-numbers API as a BYO trunk. Then point your Twilio, Telnyx, or Plivo SIP trunk at `sip.voice.x.ai`. xAI sends a signed webhook, and your server joins the call over WebSocket.

Which has better voices, Grok or ElevenLabs?

ElevenLabs offers far more choice, with 5,000+ voices and expressive Eleven v3 Conversational delivery. xAI documents 28 built-in voices plus cloning from short clips. "Better" depends on your callers and phone audio. Run a blind test at 8 kHz before deciding, because studio demos rarely match real phone quality.

Which supports more languages, Grok or ElevenLabs?

ElevenLabs covers more. Its agents support 31 languages on default models and 70+ with Eleven v3 Conversational. xAI's speech-to-speech docs list 20 language codes and claim 20+ languages, with automatic detection. For bilingual callers, test mid-call switching on both, since behavior varies by configuration.

Can I use my own LLM with Grok voice or ElevenLabs?

Not with xAI's speech-to-speech model. The Grok model is fixed, though you can set reasoning effort. ElevenLabs natively supports GPT, Claude, Gemini, and hosted Qwen models. It also accepts any OpenAI-compatible custom LLM endpoint. If model choice is a hard requirement, ElevenLabs is the only one of the two that offers it.

Do you still need to test a Grok or ElevenLabs voice agent?

Yes. Both platforms help you build and launch, but neither proves your agent handles your callers. Accents, noise, interruptions, and tool failures break agents on every stack. Independent testing with synthetic callers, such as Evalgent, measures task success and latency before real customers find the gaps.

The bottom line

xAI Grok Voice and ElevenLabs Agents both list $0.08 per minute in September 2026, but xAI bundles the model while ElevenLabs bills the LLM and telephony separately. The right choice is the one that wins on cost per successful call in a fair test on your own traffic.

Book a demo to run that side-by-side test with Evalgent: Book a demo.

Related Articles