HomeReadTools deskCartesia Sonic-3 leads real-time voice API latency benchmarks against Deepgram and ElevenLabs
Tools·Sep 6, 2026

Cartesia Sonic-3 leads real-time voice API latency benchmarks against Deepgram and ElevenLabs

An analysis of real-time Text-to-Speech API latency and pricing, evaluating Cartesia, Deepgram, ElevenLabs, PlayHT, and OpenAI to find the best engine for conversational voice agents. For…

An analysis of real-time Text-to-Speech API latency and pricing, evaluating Cartesia, Deepgram, ElevenLabs, PlayHT, and OpenAI to find the best engine for conversational voice agents.

For conversational voice bots where natural turn-taking is non-negotiable, Cartesia Sonic-3 is the clear choice. Its 85ms median Time-to-First-Byte (TTFB) makes it the only engine tested that comfortably stays below the critical 100ms threshold. Teams running high-volume, cost-sensitive voice pipelines should opt for Deepgram Aura-2, which offers the lowest bulk cost at $15.00 per million characters while maintaining a highly competitive 115ms TTFB. Skip OpenAI TTS-1 for real-time applications entirely. Its 240ms latency, driven by chunked HTTP delivery, is too slow for natural conversation.

Methodology

This review evaluates real-time Text-to-Speech (TTS) API benchmarks published in September 2026 by developer mrzitoun. The dataset aggregates median latency and pricing metrics across five primary streaming TTS providers: Cartesia (Sonic-3), Deepgram (Aura-2), ElevenLabs (Flash v2.5), PlayHT (PlayDialog), and OpenAI (TTS-1). Latency was measured as Time-to-First-Byte (TTFB) over WebSocket connections using US-East endpoints, calculated as the median across 1,000 requests.

This v0 review draws on the published claims at the dev.to blog post, the VoiceAIBench platform, and the awesome-voice-ai-latency GitHub repository. Independent verification of these benchmarks is pending. We have not independently run these 1,000-request suites or tested performance under simulated packet loss, regional routing outside US-East, or long-term session drift.

What it does

The benchmarked APIs convert text to streaming audio in real time, a critical component for WebRTC voice bots and automated phone agents. The performance profiles split the market into three distinct tiers based on latency and cost.

Cartesia Sonic-3 leads on speed

Cartesia Sonic-3 achieved a median TTFB of 85ms. This speed is vital for handling user interruptions, where the bot must instantly cease speaking and process new input. Cartesia charges $20.00 per million characters for this performance.

Deepgram Aura-2 balances cost

Deepgram Aura-2 delivers a median TTFB of 115ms at $15.00 per million characters. This makes it the most economical option for high-volume deployments that can tolerate an extra 30ms of latency to save 25% on API costs.

ElevenLabs Flash v2.5 prioritizes realism

ElevenLabs Flash v2.5 clocks in at 135ms TTFB and costs $25.00 per million characters. While slower than Cartesia and Deepgram, it remains the standard for emotional inflection, complex dialect stability, and high-fidelity voice cloning.

PlayHT and OpenAI lag behind

PlayHT PlayDialog sits at 180ms TTFB ($25.00 per million characters), while OpenAI TTS-1 trails at 240ms TTFB ($15.00 per million characters). OpenAI's reliance on chunked HTTP rather than a native streaming WebSocket protocol limits its utility in fast-paced conversational interfaces.

What's interesting and what's not

The most significant takeaway is the emergence of Cartesia Sonic-3 as the latency benchmark leader. Sub-100ms TTFB is the holy grail for voice agents because human conversational pause-to-speak latency averages around 200ms. When you add speech-to-text (STT) processing and LLM generation times, every millisecond saved on the TTS side prevents awkward, overlapping speech. Cartesia's architectural focus on raw speed over WebSockets is a genuine engineering improvement for conversational AI.

Conversely, ElevenLabs Flash v2.5 represents a classic trade-off between speed and quality. While its 135ms latency is highly usable, it pushes the limits of natural turn-taking once STT and LLM latencies are factored in. However, for applications where brand identity and emotional nuance matter more than rapid-fire interruption handling, ElevenLabs remains superior.

What is not interesting is OpenAI's positioning in this space. While TTS-1 is highly cost-competitive at $15.00 per million characters, its 240ms latency makes it a non-starter for real-time voice. OpenAI has treated TTS as a secondary feature rather than a core infrastructure product, relying on chunked HTTP instead of optimizing for WebSockets.

Pricing

Pricing snapshot as of September 2026:

  • Deepgram Aura-2: $15.00 per 1 million characters.
  • OpenAI TTS-1: $15.00 per 1 million characters.
  • Cartesia Sonic-3: $20.00 per 1 million characters.
  • ElevenLabs Flash v2.5: $25.00 per 1 million characters.
  • PlayHT PlayDialog: $25.00 per 1 million characters.

Verdict

For building highly responsive, interactive voice agents, Cartesia Sonic-3 is the top pick. Its 85ms TTFB provides the necessary headroom to keep total round-trip latency under the 200ms threshold required for natural human conversation. If your application is a high-volume outbound dialer where cost-efficiency is more critical than instant interruption handling, Deepgram Aura-2 is the correct choice. It offers a strong balance of 115ms latency and a low $15.00 per million characters price point. Avoid OpenAI TTS-1 for any conversational use case.

What we'd test next

In our next evaluation phase, we want to benchmark these APIs under real-world network degradation. We need to measure how jitter, packet loss, and regional routing (such as EU-Central and AP-Southeast endpoints) affect TTFB. Additionally, we plan to run comparative tests on how effectively each provider handles dynamic stream interruption, specifically measuring the latency between sending a WebSocket stop signal and the physical termination of the audio stream.

The investor read

The real-time voice market is experiencing a massive capital and engineering shift toward low-latency infrastructure. As LLM reasoning speeds improve, the bottleneck for voice agents has shifted entirely to the audio pipeline (STT and TTS). Cartesia's sub-100ms performance positions it as a highly attractive acquisition target or venture play, challenging ElevenLabs' market dominance. While ElevenLabs has captured the high-margin creative and narrative markets, Cartesia and Deepgram are capturing the high-volume enterprise utility market (customer service, outbound calling, real-time translation). Investors should watch if ElevenLabs' Flash model line can close the latency gap, or if specialized, speed-first infrastructure like Cartesia will monopolize the voice-agent platform layer.

Sources · how we verified
  1. Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)
  2. VoiceAIBench
  3. awesome-voice-ai-latency

Every claim ties to a primary source. See our methodology.

Reported by the Riley desk on Founderr Pulse’s Tools beat. Every factual claim is tied to a primary source and linked; anything that can’t be stood up doesn’t run. Founderr (RIKHATH LLC) is the accountable publisher and corrects in place. How we work · About · File a correction.
R
Riley

The Riley desk covers tools — what founders are building with, switching to, and abandoning. Every claim is sourced and linked. Operated by Founderr (RIKHATH LLC) See the desk →

Founderr Pulse — free & independent. The desk for people who build & back.
Cartesia Sonic-3 leads real-time voice API latency benchmarks against Deepgram and ElevenLabs · Founderr Pulse