Open door for builders.
LiveKit vs Pipecat vs Dograh: Choosing an Open-Source Voice Agent Stack for In-House Teams

On this page
Most "LiveKit vs Pipecat vs Dograh" searches come from one situation. A product or ops team at a mid-size company wants a phone agent, does not want to pay a per-minute platform tax forever, and has noticed that three open-source names keep coming up. The trouble is that these three are not the same kind of thing. Comparing them feature by feature is like comparing Postgres, an ORM and a CMS.
This guide sorts them by layer first. Then it gives you what the vendor pages do not: verified license terms (including one model license that matters commercially), who funds each project, how often each one shipped in the last six months, cost per minute at stated assumptions, the self-hosting break-even, migration paths, and one test harness that works on all three.
The site already has deep two-way comparisons. If you are only choosing between LiveKit and Pipecat, start with how to choose between LiveKit and Pipecat and the hidden gaps between LiveKit and Pipecat. This post is about where Dograh fits and how to frame the three-way choice.
All repository figures below were checked on October 4, 2026. Star counts move daily, so treat them as approximate.
What each one is, by layer
Start with the stack a phone agent needs, from the wire up:
1. Telephony ingress. PSTN numbers, SIP trunks or carrier WebSocket streams (Twilio Media Streams, Telnyx, Vonage).
2. Media transport. WebRTC rooms, RTP, or WebSocket audio frames, plus codec conversion (8 kHz G.711 on the phone side, 16 or 24 kHz for most STT and TTS models).
3. Orchestration. The real-time loop: VAD, turn detection, STT streaming, LLM calls, TTS streaming, interruption handling.
4. Conversation logic. Prompts, tools, state, multi-step flows, handoffs.
5. Application layer. A UI for non-engineers, versioning, outbound campaigns, webhooks, call records, post-call analysis, API keys, multi-tenant auth.

LiveKit: transport, SIP and an agent framework from one vendor
LiveKit is three things that ship from one company. livekit/livekit is a WebRTC media server (an SFU). livekit/sip is a separate SIP service that bridges phone calls into LiveKit rooms. livekit/agents is the Python agent framework (with a Node.js version) that joins a room as a participant and runs the STT, LLM and TTS loop.
The framework currently uses an `AgentServer` with an `@server.rtc_session()` entrypoint, and endpointing and interruption settings live under `turn_handling`, per the LiveKit server options docs. The key architectural fact: in LiveKit, a phone call becomes a SIP participant in a WebRTC room, and your agent is another participant. The media layer is LiveKit's own.
Pipecat: an orchestration pipeline with pluggable transports
Pipecat is a Python framework where audio, text and control messages flow as frames through a chain of processors. It does not own a media server. Instead it has transports: Daily WebRTC, `SmallWebRTCTransport`, a FastAPI WebSocket transport with carrier serializers (Twilio, Telnyx, Plivo, Exotel, Vonage, Genesys), and even a `LiveKitTransport`. That last one means Pipecat can run inside LiveKit rooms. We cover that pattern in using LiveKit and Pipecat together.
Recent naming matters if you copy older tutorials. In v1.3.0, `PipelineTask` became `PipelineWorker`, and `PipelineRunner` became `WorkerRunner`. The old names still work but emit a `DeprecationWarning`, per the v1.3.0 release notes.
Dograh: an application platform on a pinned Pipecat fork
Dograh describes itself as an "open source voice AI platform" and a "self-hosted alternative to Vapi and Retell". That description is the company's own positioning. What can be verified from the repository:
- It is built on Pipecat, but on a fork. The repo contains a `pipecat` git submodule pointing to `github.com/dograh-hq/pipecat`, a fork of `pipecat-ai/pipecat`. The submodule is pinned to a specific commit. Dograh 1.48.0 (October 2, 2026) includes "Upgrade Pipecat to v1.12.0", per the Dograh changelog.
- It adds the application layer. A visual workflow builder (start, agent, global, end-call, webhook, trigger and QA nodes), a browser "Test Audio" and "Test Chat" panel, outbound campaigns, call records with cost, webhooks, knowledge bases, Python and Node SDKs, and an MCP server so coding assistants can create and edit workflows.
- Telephony is built in. The README lists Twilio, Vonage, Telnyx, Plivo, Vobiz, Cloudonix and Asterisk ARI. The docs add Exotel and VICIdial (through ARI).
- It ships as a Docker Compose stack. PostgreSQL, Redis, MinIO, the API and the UI, plus nginx and coturn on remote installs.
- There is a hosted version at app.dograh.com, priced on the Dograh pricing page at a 1 cent per minute platform fee with 10 concurrent calls included.
So the closest analogy is: LiveKit is infrastructure plus a framework, Pipecat is a framework, and Dograh is an open-source Vapi-style product built with a framework. Moving up the layers buys you time to first call and costs you control at the layers below.
The numbers that matter for a long-term bet
| LiveKit | Pipecat | Dograh | |
|---|---|---|---|
| Main repo | livekit/agents (+ livekit/livekit, livekit/sip) | pipecat-ai/pipecat | dograh-hq/dograh |
| License | Apache-2.0 (code); turn-detector models under LiveKit Model License | BSD-2-Clause; Smart Turn model repo also BSD-2-Clause | BSD-2-Clause |
| Stars (approx.) | agents 14.4k, server 21.3k (Oct 4, 2026) | 16.2k | 5.5k |
| Repo created | agents: Oct 2023 | Dec 2023 | Sep 2025 |
| Latest release | livekit-agents 1.8.4 (Oct 1, 2026) | 1.12.0 (Sep 25, 2026) | 1.48.0 (Oct 2, 2026) |
| Maintainer | LiveKit Inc | Daily ("Maintained by Daily and the community") | Zansat Technologies Private Limited (per README) |
| Disclosed backing | Series C, $100M at $1B valuation, Jan 2026 | Daily: $40M Series B, Nov 2021 | No funding round found in public sources |
| Hosted option | LiveKit Cloud | Pipecat Cloud (Daily) | Dograh Cloud |
Sources: GitHub repositories and API for counts and licenses; LiveKit Series C post, led by Index Ventures with Salesforce Ventures, Hanabi Capital, Altimeter and Redpoint; Daily's Series B post, led by Renegade Partners. Dograh's README says it was "founded by YC alumni and exit founders". That is the company's own statement, and it names no investors.
Contributor concentration
Stars measure attention, not resilience. A better signal for a production bet is how concentrated the commit history is. Using GitHub's top-30 contributor lists on October 4, 2026:
- Dograh: the top two accounts have 752 of 846 contributions, about 89%.
- Pipecat: the top two have 8,058 of 12,340, about 65%.
- LiveKit Agents: the top two have 1,593 of 3,109, about 51%.
Every young project looks like this. Two founders writing most of the code is normal at one year old. But it is a fact worth pricing in: if Dograh's core team changes direction, you own a fork. Because the license is BSD-2, you are allowed to, and the platform is still built on Pipecat, which has its own broader maintainer base.
Release cadence and breaking changes
Counting releases from April 4 to October 4, 2026:
- LiveKit Agents: 35 stable releases of the core Python package (1.5.2 to 1.8.4), plus 4 release candidates. Plugin packages are counted separately and are not in this number. Changes you would feel: 1.6.8 deprecated the `console` and `dev` modes in favor of the `lk agent` CLI, and 1.7.0 renamed sensitive trace attributes and log fields. If your dashboards query those names, they break. The `turn_handling` consolidation landed in 1.5.0 (March 2026). The old keyword arguments are deprecated and scheduled for removal in 2.0.
- Pipecat: 15 releases (1.0.0 to 1.12.0). The window opened with 1.0.0, which removed `OpenAILLMContext` (use `LLMContext`), removed flat imports like `from pipecat.services.openai import OpenAILLMService`, and moved VAD and turn analyzers out of `TransportParams`. Then 1.3.0 renamed `PipelineTask` and `PipelineRunner`. And 1.10.0 removed `LmntTTSService` because LMNT shut down.
- Dograh: 30 releases (1.21.0 to 1.48.0), roughly one a week. Releases that would affect a self-hoster include 1.29.0 (multi-worker deployment, required for the scaling layout in the docs), 1.47.0 ("let external-turn STT decide the turn start, retire provisional_vad", which changes turn behavior), and 1.48.0 (Pipecat 1.12.0 upgrade and a new default Gemini Live model).
Here is what a frequent release cadence means in practice. On LiveKit or Pipecat, you choose when to absorb a framework change, and you read one changelog. On Dograh, you get Pipecat changes bundled into Dograh releases. Someone else does the upgrade work, but you are taking two projects' behavior changes at once, on Dograh's schedule. A default model swap or a turn-detection change inside a platform release can move your latency and interruption numbers even if you changed nothing. That is a regression-testing problem, not a reason to avoid the platform.
Licenses and what they mean commercially
All three codebases use permissive licenses. Apache-2.0 (LiveKit) and BSD-2-Clause (Pipecat and Dograh) both allow commercial use, modification, private forks and embedding in proprietary products, without any obligation to publish your changes. Neither is copyleft. Neither has the network-use clause that the AGPL has, so running a modified version as a hosted service does not trigger disclosure.
The practical differences are smaller, but they are real:
| Question | Apache-2.0 (LiveKit) | BSD-2-Clause (Pipecat, Dograh) |
|---|---|---|
| Can we fork privately and sell a product on it? | Yes | Yes |
| Must we publish modifications? | No | No |
| Attribution required? | Yes, keep LICENSE and NOTICE files | Yes, keep copyright notice and license text |
| Must modified files be marked as changed? | Yes | No |
| Explicit patent grant from contributors? | Yes, with patent-retaliation clause | No explicit patent grant |
| Trademark use? | Not granted | Not addressed; assume not granted |
The clause that does affect the decision is not in the code licenses. LiveKit's turn-detector models are under a separate LiveKit Model License. It lets you use the models freely, but only with the LiveKit Agents framework, not standalone and not with other frameworks. It also bars you from using the models or their outputs to improve other models.
In practice, if a big part of your endpointing quality comes from LiveKit's turn detector, that quality does not travel with you if you migrate to Pipecat or Dograh. Pipecat's Smart Turn model repository is BSD-2-Clause, so it has no such restriction. Dograh uses whatever turn handling its pinned Pipecat version provides, plus its own changes.
A second governance point: none of the three is under a neutral foundation. Each is steered by one company that also sells a hosted version. That is common in open-source infrastructure. It means roadmap priorities will follow the hosted product. One example is in LiveKit's own docs: enhanced noise cancellation is offered as a LiveKit Cloud feature, so a self-hosted LiveKit deployment needs its own noise suppression.
Telephony: three different ways audio reaches your agent
How a phone call becomes audio frames is the biggest hidden difference between the three, and it drives both cost and failure modes. The full two-way breakdown is in LiveKit vs Pipecat for telephony. The three-way summary follows.
LiveKit: SIP into a room. A SIP trunk (Twilio Elastic SIP, Telnyx and others) sends INVITEs to the LiveKit SIP service, which creates a SIP participant in a room. Self-hosting means running `livekit-server`, `livekit-sip` and Redis. Per the livekit/sip README, "the SIP service and the LiveKit server communicate over Redis" and the SIP service needs a public IP. LiveKit's self-hosting docs also require port 5060 and UDP 10000 to 20000 to be reachable. RTP arrives as G.711 at 8 kHz and is transcoded into the room's codec.
Pipecat: you choose. The common path is a carrier WebSocket. Twilio's `
Dograh: carrier WebSockets and Asterisk. Dograh's telephony webhook docs list the exact formats: Twilio, Plivo and Vobiz send 8 kHz μ-law, base64-encoded in JSON messages. Vonage sends 16 kHz linear PCM in binary frames. Asterisk ARI sends 8 kHz linear PCM through `externalMedia`. Every stream lands on `/api/v1/telephony/ws/{workflow_id}/{organization_id}/{workflow_run_id}`.
Two details in those docs are worth knowing before a security review:
- When `TELEPHONY_WS_TOKEN_SECRET` is set, Dograh adds an HMAC token as a fourth path segment, not a query parameter. The reason given is that carriers do not reliably forward query strings, and Twilio strips them from `
` entirely. If you put an auth token in a Twilio stream URL query string on any stack, it will silently disappear. - Twilio webhook signatures cover the full URL including the query string, so a proxy that rewrites or reorders parameters will break verification.
The cost consequence is concrete, and the next section puts numbers on it. A Twilio Media Streams path pays Twilio's voice rate plus a Media Streams fee. A SIP trunk path pays only the trunk rate. Both Pipecat and Dograh can avoid Media Streams: Pipecat through Daily or LiveKit SIP, Dograh through Asterisk ARI with your own trunk, or Dograh's managed SIP on its cloud. But the quick-start path for both is the carrier WebSocket.
Cost per minute at stated assumptions
Every number in this section is either a published list price, linked, or a labeled assumption. Swap in your own contract rates.
Shared provider stack (assumption). To compare the stacks rather than the models, every option uses the same providers. These are the example per-minute prices on the LiveKit pricing page: Deepgram Nova-3 STT at $0.0048, Cartesia Sonic 3 TTS at $0.03 and GPT-4.1 mini at $0.0015. That totals $0.0363 per minute. Direct provider contracts will differ. This line is identical across stacks, so it does not change the ranking.
Telephony list prices (US, inbound local).
- Twilio voice: $0.0085/min, plus Media Streams at $0.0044/min (Twilio voice pricing).
- Twilio Elastic SIP origination: $0.0034/min (Twilio SIP trunking pricing).
- LiveKit Cloud US local number: $0.01/min inbound.
Platform fees.
- LiveKit Cloud agent session minutes: $0.01 overage on the Ship and Scale plans.
- Pipecat Cloud agent-1x: $0.01 per active minute (Pipecat Cloud pricing).
- Dograh Cloud: $0.01 per minute platform fee.
Self-hosted compute (assumption). A 4 vCPU, 8 GB VM at an illustrative $0.17/hour. This is Dograh's documented minimum for a remote install. Assume 20 concurrent calls per VM at 60% average utilization. That gives $0.17 ÷ 60 ÷ (20 × 0.6) = $0.00024/min, rounded to $0.0003 to cover Postgres, Redis and storage. You must load-test your own concurrency per VM. None of the three projects publishes a calls-per-vCPU figure.
| Option | Providers | Telephony | Platform or compute | Total per minute |
|---|---|---|---|---|
| A. LiveKit Cloud, LiveKit number | $0.0363 | $0.0100 | $0.0100 | $0.0563 |
| B. LiveKit self-hosted, Twilio SIP trunk | $0.0363 | $0.0034 | $0.0003 | $0.0400 + ops |
| C. Pipecat Cloud, Twilio Media Streams | $0.0363 | $0.0129 | $0.0100 | $0.0592 |
| D. Pipecat self-hosted, Twilio Media Streams | $0.0363 | $0.0129 | $0.0003 | $0.0495 + ops |
| E. Dograh Cloud (BYOK), Twilio Media Streams | $0.0363 | $0.0129 | $0.0100 | $0.0592 |
| F. Dograh self-hosted, Asterisk + SIP trunk | $0.0363 | $0.0034 | $0.0003 | $0.0400 + ops |

Three things in this table are not obvious:
1. Providers dominate. Model costs are 61% to 91% of every row. Picking a cheaper TTS moves your bill more than picking a framework.
2. The telephony ingress gap ($0.0095/min) is about the same size as any platform fee ($0.01/min). Moving from Media Streams to a SIP trunk saves roughly what self-hosting saves. Many teams self-host to dodge a platform fee and keep paying the Media Streams premium.
3. The three hosted options cost nearly the same per minute. Choose a hosted tier on concurrency limits, regions and support, not on the headline rate. Dograh Cloud includes 10 concurrent calls. LiveKit's Ship plan lists 20 concurrent agent sessions. Pipecat Cloud lists concurrency as unlimited.
The self-hosting break-even
Self-hosting is not free. Someone patches the VM, rotates TURN secrets, upgrades Postgres, watches WebSocket memory and absorbs the 30-plus releases each project ships every six months. Assume that costs 0.25 of an engineer at a $180,000 loaded annual cost. That is $3,750 per month (assumption).
Break-even minutes per month = ops cost ÷ (platform fee − compute per minute)
= $3,750 ÷ ($0.0100 − $0.0003) = about 387,000 minutes per month
At an average three-minute call, that is about 129,000 calls a month, or roughly 4,300 a day. Below that volume, a hosted tier is cheaper once you count people. Above it, self-hosting wins, and the gap grows with volume. The real reasons to self-host below break-even are data residency, on-prem requirements and control, not cost. Our voice agent total cost of ownership guide walks through the people-cost side in more detail.

What self-hosting each one involves
LiveKit self-hosted. For phone calls you run three services (server, SIP, Redis), plus your agent servers. The media server is Go and scales horizontally with Redis. The SIP service needs a public IP and a wide UDP range for RTP. Agent servers are separate processes that register with the server and get jobs dispatched to them. The most moving parts, but each one is purpose-built.
Pipecat self-hosted. For a Twilio WebSocket deployment, you run your own FastAPI app with one pipeline per call. There is no bundled media server, database or UI. You decide where call records go and how to balance load. This is the leanest footprint and the most of your own code. Pipecat's own issue tracker has asked for battle-tested self-hosted blueprints, which tells you the deployment layer is yours to design.
Dograh self-hosted. One script brings up the stack. The Dograh scaling docs explain the runtime clearly, and three details matter for capacity planning:
- The API container runs the call pipelines. It starts `FASTAPI_WORKERS` separate uvicorn processes, each on its own port, and nginx balances them with `least_conn`. The docs explain why they avoided `uvicorn --workers`. With pre-fork, the kernel spreads new TCP connections at `accept()`, and long-lived WebSockets stick to whichever worker accepted them, so "a handful of unlucky workers end up handling most of the streaming traffic". Remember this if you scale any WebSocket-based voice stack, Pipecat included.
- Some components are singletons by design. `ari_manager` and `campaign_orchestrator` do not scale with workers. Postgres, Redis and MinIO each run as a single container. For production, the docs say to point `DATABASE_URL` at a managed database.
- Budget 300 to 500 MB of RAM per worker, on a minimum of 4 vCPUs and 8 GB RAM.
Two smaller gotchas. Anonymous telemetry is on by default (set `ENABLE_TELEMETRY=false`). And if you build from source, you must run `git submodule update --init --recursive` after pulling, "or the Docker build step will not pick up pipecat changes". That warning is in the Docker deployment docs. Forgetting it means you are running a different Pipecat than you think.
A decision matrix for in-house teams
This matrix is built for a 50 to 500 person company whose own engineers run the agent. Read across a row and note which column describes you. The column you land in most often is your starting point.
| Criterion | Favors LiveKit | Favors Pipecat | Favors Dograh |
|---|---|---|---|
| Voice team size | 2+ engineers who will own media infra | 1 to 3 Python engineers who want full control | 1 engineer plus ops or product people who edit flows |
| Custom audio logic (custom VAD, mixing, parallel pipelines, video) | Strong: room model, multi-participant | Strongest: any processor anywhere in the frame chain | Weak: you would be editing a fork of a fork |
| Telephony needs | SIP trunks, dispatch rules, transfers in one vendor | Carrier WebSockets now, SIP via Daily or LiveKit later | Outbound campaigns, many carriers, Asterisk or VICIdial shops |
| Ops capacity | Can run Redis, an SFU and SIP, or pay for Cloud | Can run a FastAPI fleet and design the deployment | Wants one Compose stack or a hosted tier |
| License and governance | Apache-2.0 code, but turn models are LiveKit-only | BSD-2, including the Smart Turn model | BSD-2, younger project, concentrated maintainers |
| Vendor dependency | One vendor for media, SIP and framework | Daily maintains it, but transports are swappable | Dograh team plus an upstream Pipecat dependency |
| Time to first call | Fast with Cloud and `lk agent`; slower self-hosted | Fast with `pipecat init` for one bot; slower to build product features | Fastest: browser test with bundled keys after a 2 to 3 minute start |
| Debugging depth | OpenTelemetry traces, built-in test helpers and judges | Full code access, Whisker debugger, OpenTelemetry | Run records, tracing, an OTEL endpoint (1.48), BigQuery call event export; deep issues mean reading Dograh and Pipecat code |
A few patterns come out of this:
- Non-engineers need to change flows weekly: Dograh, or Pipecat Flows plus a UI you build.
- The agent is part of your product (voice inside vertical SaaS, custom audio handling): Pipecat or LiveKit, because you will need code paths a visual builder does not expose.
- The main need is telephony at scale with transfers and SIP: LiveKit, or Pipecat on LiveKit transport.
- You are leaving Vapi or Retell and want the same mental model on your own infrastructure: Dograh is the closest match. Note that this matches Dograh's own positioning. Verify it on your call flows before you commit. LiveKit vs Vapi covers the build-versus-buy side.
What each stack makes hard to test
Every stack has a cheap testing path that skips something important. Knowing which part each one skips tells you where your blind spots are.
LiveKit
LiveKit has the best built-in testing of the three. Per the LiveKit unit test docs, you can run text turns through an `AgentSession` and assert on events. Simplified example:
# Simplified: LiveKit Agents text-mode test (pytest + pytest-asyncio)
import pytest
from livekit.agents import AgentSession
from livekit.plugins import openai
from my_agent import SchedulingAgent # your Agent subclass
@pytest.mark.asyncio
async def test_reschedule_asks_for_slot():
async with (
openai.LLM(model="gpt-4.1-mini") as llm,
AgentSession(llm=llm) as session,
):
await session.start(SchedulingAgent())
result = await session.run(user_input="I need to move my appointment to Friday")
await (
result.expect.next_event()
.is_message(role="assistant")
.judge(llm, intent="Confirms which appointment or offers Friday time slots")
)What it skips: STT errors on 8 kHz audio, the turn detector, VAD, TTS timing and the SIP leg. These tests run on text. They catch prompt and tool regressions, not "the agent cut off a caller who paused to find their card number". Also note that the same LLM can serve as agent and judge in this pattern. Use a different judge model for anything you gate releases on. Our guide to the limits of LLM-as-judge explains why. The full LiveKit workflow is in our LiveKit voice agent testing guide.
Pipecat
Pipecat's frame model makes unit-testing a single processor easy: push frames in, assert on frames out. The hard parts are races between frames. Examples are an interruption frame arriving while TTS audio is still queued, or context updates after an interruption that leave the LLM believing it said text the caller never heard.
The serializer is a second blind spot. A pipeline tested over WebRTC at 16 kHz runs a different code path from the same pipeline behind `TwilioFrameSerializer` at 8 kHz μ-law. Test on the transport you ship. See the Pipecat voice agent testing guide.
Dograh
Dograh's Test Chat lets you edit or replay user turns and regenerate replies and node transitions. That makes prompt iteration fast, but it is text-only. Four platform behaviors need deliberate testing:
- Transitions are decided by an LLM. Per the Dograh workflow docs, each edge has "a natural language description of when to move on" that "the LLM evaluates". The same caller utterance can take different paths on different runs. You need repeated runs per scenario and a pass rate, not one green check.
- The QA node reads only the transcript. The QA node runs an LLM review after the call, with default tags like `DEAD_AIR` and `USER_FRUSTRATED`, a 1 to 10 score, and settings such as a 15-second minimum duration and a sample rate. It is useful for triage. But it cannot hear latency, talk-over, TTS mispronunciations or audio glitches. If "Use Workflow's LLM" is on, the model grades its own work. We explain the gap in transcript vs audio evaluation.
- Platform upgrades change runtime behavior. As covered above, a release can change the Pipecat version, turn handling or default models. Pin the Dograh image version in production and run your regression suite before every upgrade.
- The browser test path is not the phone path. Test Audio uses WebRTC through the browser, with coturn on remote installs. Phone calls arrive as 8 kHz μ-law over carrier WebSockets. Latency and STT accuracy differ between them.
One evaluation harness for all three
The only interface all three stacks share is a phone number. That makes a PSTN black-box harness the fairest way to compare them, and it doubles as your regression suite after you choose. The design:
1. A test number calls the agent number (or the agent calls the test number, for outbound flows).
2. The caller side plays scripted audio: utterances, deliberate pauses, barge-ins and noise mixed at fixed SNRs.
3. The call is recorded dual-channel, so caller and agent audio are on separate channels.
4. Metrics are computed from audio timing plus the structured outcome each stack reports (webhook payloads, tool-call logs).
Simplified Python for the call and the timing analysis:
# Illustrative: framework-neutral PSTN harness (Twilio REST + numpy)
import numpy as np, soundfile as sf
from twilio.rest import Client
client = Client(ACCOUNT_SID, AUTH_TOKEN)
def place_test_call(agent_number: str, script_url: str) -> str:
# The caller side plays a scripted WAV (utterances + pauses + barge-ins).
twiml = f"<Response><Play>{script_url}</Play><Pause length='20'/></Response>"
call = client.calls.create(
to=agent_number, from_=TEST_NUMBER, twiml=twiml,
record=True, recording_channels="dual",
)
return call.sid
def speech_mask(x, sr, frame_ms=20, thresh_db=-35):
n = int(sr * frame_ms / 1000)
frames = x[: len(x) // n * n].reshape(-1, n)
rms_db = 20 * np.log10(np.sqrt((frames ** 2).mean(axis=1)) + 1e-9)
return rms_db > thresh_db # True = speech in this 20 ms frame
def response_gaps(wav_path, frame_ms=20, min_silence_ms=300):
audio, sr = sf.read(wav_path) # dual-channel recording
caller, agent = speech_mask(audio[:, 0], sr), speech_mask(audio[:, 1], sr)
gaps, f = [], frame_ms
for i in range(1, len(caller)):
if caller[i - 1] and not caller[i]: # caller stopped
j = i
while j < len(agent) and not agent[j]:
j += 1
if (j - i) * f >= min_silence_ms and j < len(agent):
gaps.append((j - i) * f) # ms until agent audio
return np.percentile(gaps, [50, 95]) if gaps else NoneCheck which channel holds which leg in your carrier's recording before trusting the numbers. The energy threshold also needs tuning against your noise conditions. The point is that nothing here knows which framework answered the call.
For outcomes, use each stack's native record. On Dograh, you can start outbound test runs through the documented endpoint `POST /api/v1/public/agent/workflow/{workflow_uuid}` with an `X-API-Key` header and a JSON body containing `phone_number` and an optional `initial_context`. The response returns a `workflow_run_id`, which you use to fetch the transcript, recording and gathered data. On LiveKit and Pipecat, emit the same fields from your own tool handlers and end-of-call hooks, so all three produce one comparable outcome record per call.
How to run a fair LiveKit vs Pipecat vs Dograh bake-off
1. Fix everything except the stack. Use the same STT, LLM, TTS, voice, prompt text and tools on all three. If one stack forces a different provider, record it as a confound.
2. Fix the telephony path. Run all three through the same carrier and ingress type. Either all on Twilio Media Streams, or all on SIP trunks. Mixing them makes the comparison about Twilio, not the frameworks.
3. Write 40 scenarios that match your real traffic: 25 happy paths, 8 edge cases (spelled names, account numbers, corrections), 4 barge-in scripts and 3 off-topic or adversarial callers. Give each scenario a machine-checkable expected outcome, such as "appointment moved to 2026-10-09 10:00, confirmation spoken".
4. Cross scenarios with 3 conditions: clean audio, café noise at 10 dB SNR and café noise at 5 dB SNR. Mix the noise into the caller script offline so it is identical across stacks.
5. Run each cell 3 times, which is 360 calls per stack. Repeats matter most for Dograh, because of LLM-decided transitions, but every stack's LLM output varies.
6. Score with fixed definitions. Response gap p50 and p95 (caller speech end to agent audio start). Barge-in stop time (caller speech start to agent audio stop). False cutoff rate (the agent starts speaking while the caller's scripted utterance is still playing). Task success (outcome record matches the expected outcome). Report each with its variance across the 3 repeats.
7. Size the sample for the difference you care about. To detect task success of 85% vs 92% at 95% confidence and 80% power, use n = (1.96 + 0.84)² × (0.85 × 0.15 + 0.92 × 0.08) ÷ 0.07² = 7.84 × 0.2011 ÷ 0.0049 ≈ 322 calls per stack. The 360-call design clears that. If you expect a smaller gap, the sample size grows with the inverse square of the gap.
8. Re-run the suite after every upgrade. That means each Dograh image bump, each LiveKit Agents minor release and each Pipecat release you adopt. Store results by version so a regression points straight at a release.
Migration paths between the three
None of these choices is permanent. Here is what moving actually involves.
Dograh to Pipecat. This is the shortest path, because Dograh already runs on Pipecat. Provider choices map one to one onto Pipecat services. Export each agent's graph through the Get Agent API, which returns the full workflow definition. Then rebuild it in Pipecat Flows, which uses a similar node-and-transition model. What you rebuild yourself: campaigns, the UI, run storage, webhooks and the QA step.
Pipecat to LiveKit, gradually. Move Pipecat onto `LiveKitTransport` first, so calls arrive through LiveKit SIP and rooms while your pipeline code stays the same. Then port the conversation logic to LiveKit Agents if you want the built-in turn detector and test helpers. The turn detector only works inside LiveKit Agents, so the payoff only arrives at the second step.
LiveKit to Pipecat. Keep LiveKit as the media and SIP layer, and swap the orchestration to Pipecat with `LiveKitTransport`. Expect to replace the LiveKit turn-detector models. Their license does not allow use outside LiveKit Agents.
Vapi or Retell to Dograh. The concepts carry over (graph of prompt nodes, tools, webhooks, outbound campaigns), but there is no importer, so you rebuild flows by hand or through the API and MCP server. Our voice agent vendor migration guide covers running the old and new stacks in parallel before cutover. The lock-in tradeoffs are in voice AI vendor lock-in.
Any stack to any stack. Keep three things stack-neutral from day one: your scenario suite, your outcome schema and your phone-level harness. They are the only parts of a voice agent that survive a migration unchanged. They are also how you prove the new stack is not worse.
Testing your choice before you commit
Do not decide from feature tables, this one included. Build the same agent, the scheduling flow or support triage you actually need, on the two stacks that won your decision matrix. Run the 40-scenario, 3-condition suite above through the same carrier. Compare p95 response gap, false cutoffs and task success with their variance across repeats.
That is usually one to two weeks of work for an in-house team. Most of the time goes into the harness, not the agents. Independent evaluation helps in two places. In a pre-commit bake-off, the scoring is done by someone with no stake in which framework wins. After launch, the same suite runs as a regression gate on every framework or platform upgrade, which matters most when your stack absorbs a release every week. Evalgent does both: it places real calls against your LiveKit, Pipecat or Dograh agent, under controlled noise and interruption conditions, and scores them from the audio, not just the transcript.
Frequently asked questions
Is Dograh built on Pipecat?
Yes. The Dograh repository includes a `pipecat` git submodule that points to `dograh-hq/pipecat`, a fork of `pipecat-ai/pipecat`, pinned to a specific commit. Dograh periodically upgrades it. Version 1.48.0 moved to Pipecat 1.12.0. Pipecat changes reach Dograh users only when Dograh bumps the submodule.
What license does Dograh use?
Dograh is licensed under BSD-2-Clause, the same license as Pipecat. You can use it commercially, modify it, keep a private fork and embed it in proprietary products, as long as you keep the copyright notice and license text. There is no copyleft or network-use disclosure requirement.
Is Dograh a good open-source Vapi alternative?
It is the closest of the three to Vapi's model: a visual flow builder, built-in telephony, campaigns, webhooks and a hosted tier. That positioning is Dograh's own. The project started in September 2025 and two maintainers hold most of the commits, so validate it on your call flows and plan for upgrade testing.
Which is cheapest per minute: LiveKit, Pipecat or Dograh?
At list prices, the three hosted tiers cost nearly the same, each adding about $0.01 per minute. Your STT, LLM and TTS choices and your telephony ingress move the total far more. A SIP trunk costs about $0.0095 per minute less than Twilio Media Streams. Self-hosting pays off only at high volume.
When does self-hosting beat a hosted tier?
At the stated assumptions (a quarter of an engineer at $180,000 loaded per year, $0.0003 per minute compute), self-hosting breaks even near 387,000 minutes per month, or about 129,000 three-minute calls. Below that, choose self-hosting for data residency or control, not for cost.
Can I use LiveKit's turn detector with Pipecat or Dograh?
No. LiveKit's turn-detector models are under the LiveKit Model License, which allows use only with the LiveKit Agents framework and bars standalone use or use with other frameworks. Pipecat's Smart Turn model is BSD-2-Clause and has no such restriction.
How often do these projects ship breaking changes?
From April to October 2026, LiveKit Agents shipped 35 stable releases, Pipecat 15 and Dograh 30. Pipecat 1.0.0 and 1.3.0 removed or renamed core APIs. LiveKit deprecated its console and dev modes and renamed trace fields. Dograh changed turn handling and upgraded Pipecat inside regular releases.
How do I compare all three stacks fairly?
Use the same providers, prompts and carrier path on each. Then call each agent over the phone with scripted audio, record both channels, and score response gap, barge-in stop time, false cutoffs and task success. About 322 calls per stack can detect an 85% versus 92% success difference.
The bottom line
LiveKit, Pipecat and Dograh are not three versions of one product: LiveKit owns the media and SIP layer, Pipecat owns the orchestration pipeline, and Dograh adds a Vapi-style application layer on a pinned Pipecat fork. Pick the highest layer that still gives you the control your call flows need, then prove the choice with the same phone-level test suite you will keep running on every upgrade.
Related Articles

How to automate voice agent testing: synthetic callers vs manual QA
Learn how ai test automation replaces manual QA for voice agents. Compare synthetic callers vs human testers, with a 5-step framework to scale without hiring.
Read more
AI Agent Testing vs Voice Agent Testing: What General Tools Miss for Voice
AI agent testing measures text outputs. Voice agent testing measures behaviour through an acoustic pipeline. Five failure categories general tools miss.
Read more