Best Low Latency Text to Speech APIs for Voice Agents 2026

Compare eight low-latency TTS APIs for voice agents by TTFA, streaming, interruptions, pricing, and the production decision each changes.

Wednesday, August 26, 2026Omid Saffari
Tools
  • DDeepgram Flux TTS
  • RRime Mist v3
  • IInworld Realtime TTS-2 Flash
  • EElevenLabs Flash v2.5
  • CCartesia Sonic-3.6
  • HHume Octave 2
  • GGradium
  • OpenAI GPT-4o Mini TTS
  • PPika Speech
  • ElevenLabs
Best Low Latency Text to Speech APIs for Voice Agents 2026

Deepgram Flux TTS is the best overall low-latency API for a voice agent in 2026 because its published as-low-as-80-ms first audio comes with native barge-in state, not just fast synthesis. Pika can generate a minute of 48 kHz speech in about 1.2 seconds, but its live API is asynchronous and explicitly not for real-time streaming, so that headline does not buy a faster phone call.

The Short Answer

Choose Deepgram Flux TTS for an English-language voice agent where callers interrupt, change their minds, and expect the agent to remember exactly what it finished saying. Its latency claim is not the smallest number on this page. Its advantage is that the streaming protocol reports both the text the caller heard and the text that was cut off, which protects the conversation state after barge-in.

Choose Rime Mist v3 when transparent percentile reporting and very fast first audio matter most. Choose Inworld Realtime TTS-2 Flash when multilingual reach and character cost dominate. Choose ElevenLabs Flash v2.5 when a mature multilingual voice ecosystem matters enough to accept a higher character rate and a text-normalization chore.

Prices below were verified on vendor pages on August 26, 2026. “Starting price” means the smallest public entry point, not the likely production bill.

ToolBest forStarting priceFree trial
Deepgram Flux TTSBarge-in-heavy English voice agentsFree through Sep. 12, then $0.045/1K charsYes, $200 credit
Rime Mist v3Transparent low-latency benchmarking$0.03/1K charsYes, amount conflicts on page
Inworld TTS-2 FlashMultilingual scale and low character costStart free, then $15/1M charsYes, up to 70 min
ElevenLabs Flash v2.5Mature multilingual voice operationsFree/PAYG, then $0.05/1K charsYes, 20K chars
Cartesia Sonic-3.6Persistent WebSocket contexts$0, about 27 minYes
Hume Octave 2Expressive conversational delivery$0, about 10 minYes
GradiumOne regional TTS and STT vendor$0, about 1 hour, noncommercialYes
OpenAI GPT-4o Mini TTSExisting OpenAI stacks$0.60/M text tokens plus $12/M audio tokensNo

The explicit decision rule is simple: interruption correctness beats a lower isolated latency number. Once two candidates handle your barge-in path correctly, compare p95 first audio through the real telephony route, then compare effective cost and voice fit. A vendor can win a server-side race and still lose the call after network transit, buffering, audio conversion, or a broken interruption.

If you need a packaged platform instead of a component API, start with the voice-agent platform comparison. If narration, accessibility, or general audio is the job, the broader text-to-speech guide uses a different decision order.

How These APIs Were Picked

This is a verified comparison, not a claim that eight APIs were exercised in one controlled lab. Every model name, public price, tier, limit, transport, and published latency number came from a live first-party page checked on August 26, 2026. The ranking then applies a voice-agent production lens to those facts.

That distinction matters because the published latency numbers do not measure the same thing. Inworld reports a server-side P90 time to first audio byte that excludes the network. ElevenLabs publishes approximate model latency that excludes application and network delay. Rime publishes P50 and P90 time to first audio at stated concurrency, then separately estimates regional network round trip. Deepgram says “as low as” 80 ms. Hume distinguishes roughly 100 ms of model latency from roughly 200 ms until first audio in instant mode.

Those numbers are useful inside each vendor's system. They are not a clean cross-vendor benchmark.

Five criteria decided the order:

  1. First playable audio: The first-audio budget starts when the agent has text ready and ends when the caller can hear speech. Full-file generation time is a different metric.
  2. Interruption state: A real call needs to know what was spoken, what was not, and what the language model should retain after the caller barges in.
  3. Streaming transport: Persistent WebSockets, progressive input, cancellation, and usable audio formats can remove more delay than a small model-level advantage.
  4. Production economics: The relevant bill is generated characters or tokens at expected volume, plus minimums and tier effects, not the cheapest logo on a pricing page.
  5. Named walls: Language coverage, number normalization, concurrency, preview status, output encoding, commercial-use rights, and unclear benchmark definitions all change the practical winner.

Voice quality still matters, but it comes after the hard failures. A beautiful voice that starts late, says an account number ambiguously, or continues talking over the caller is not a good agent voice.

A four-way decision flow routing voice-agent requirements to Deepgram, Rime, Inworld, and Hume
Route the hard constraint first, then compare voices inside the surviving shortlist.

1. Deepgram Flux TTS: Best Overall for Voice Agents

Deepgram Flux TTS is the strongest overall choice because its voice-agent protocol treats interruption as state, not merely a stop button.

Deepgram Flux TTS product page
Deepgram Flux TTS

Deepgram's GA record publishes first audio as low as 80 ms, but the more valuable production detail arrived with general availability on August 12, 2026. On the live WebSocket, an Interrupt message cancels the current turn. The resulting SpeechInterrupted event returns text_spoken and text_remaining. That lets the agent update its memory from what the caller actually heard instead of estimating from player timing.

Consider an account-balance agent that begins, “Your balance is two thousand eight hundred...” before the caller interrupts with a payment question. A generic cancel command can stop the waveform, but the language model may still believe the entire balance was delivered. Flux gives the application the boundary needed to repair that state. The business consequence is fewer repeated facts, fewer contradictory follow-ups, and less custom playback bookkeeping.

Flux also carries conversation context across turns through its v2 Speak WebSocket. That is useful when delivery should stay calm after a complaint or become concise after repeated interruptions. It is a better fit for a multi-turn call than a narration model that sees each line as an isolated clip.

The main wall is scope. The GA catalog is 36 English voices across seven accents. If the launch requires broad multilingual coverage, Inworld or ElevenLabs moves ahead. Flux also streams raw linear16, mulaw, or alaw audio in the default Voice Agent integration, with no container or bit rate. Those formats suit telephony, but a team expecting MP3, Opus, FLAC, AAC, or WAV from the default path must change its audio assumptions or choose an Aura model.

Best for: English customer-service, scheduling, qualification, and support calls with frequent barge-in.
Standout: Native spoken-versus-unspoken text reporting after an interruption.
Pricing: Flux is free through September 12, 2026; standard Pay As You Go is $0.0450 per 1,000 characters and Growth is $0.0405 per 1,000 beginning September 13.
Free trial: Yes. Pay As You Go includes a free $200 credit with no minimum and no expiration.

Deepgram pricing, every current tier

The live Deepgram pricing page has three commercial routes. Pay As You Go begins with the $200 credit and supports up to 45 TTS REST or WebSocket connections. Growth starts at $4,000 per year in prepaid credits and raises that TTS concurrency ceiling to 60. Enterprise is custom for higher volume, private deployment, data, or support requirements.

Model pricing matters too. After the Flux promotion, Pay As You Go costs $0.0450 per 1,000 characters and Growth costs $0.0405. Aura-2 costs $0.030 and $0.027 respectively. Aura-1 costs $0.0150 and $0.0135. The cheaper Aura choices can make sense for less interactive audio, but they do not carry Flux's conversation-native interruption behavior.

The upside
What it does well
4 points

  • Native barge-in events report exactly what was spoken and what remained.
  • Conversation context carries across turns instead of resetting every line.
  • Published first audio is as low as 80 ms.
  • Cloud, VPC, on-prem, and self-hosted routes can support stricter deployment needs.
The downside
Where it falls short
3 points

  • Flux is English-only today.
  • The default voice-agent path accepts only raw linear16, mulaw, and alaw output.
  • Standard Flux pricing is higher than Rime Mist and Inworld Flash at the public character rates.

A practical Flux setup

  1. Open the live transport

    Connect the agent to the v2 Speak WebSocket and keep that connection alive through the conversation. Do not turn every sentence into a fresh batch request.

  2. Choose the phone-ready encoding

    Select linear16, mulaw, or alaw to match the downstream audio path. Avoid a conversion step in the hot path when the telephony provider already accepts one of those formats.

  3. Stream text in speakable chunks

    Send complete clauses as the language model produces them. A clause gives the voice enough context to sound natural without waiting for an entire paragraph.

  4. Treat interruption as memory

    When the caller barges in, send Interrupt. Replace the planned assistant turn with text_spoken in conversation state and discard text_remaining before the next language-model response.

  5. Measure what the caller hears

    Log text-ready time, first byte received, first frame queued, and first frame played. Optimize the slowest boundary at p95 instead of celebrating the vendor's best-case number.

Verdict: Deepgram wins this ranking because the API helps preserve conversational truth after a barge-in. Pick something else when multilingual coverage is non-negotiable or when character price is the dominant constraint.

2. Rime Mist v3: Best for Transparent Low-Latency Benchmarking

Rime Mist v3 earns second place by publishing the most decision-useful latency breakdown in this group.

Rime Mist v3 and Coda pricing and comparison page
Rime Mist v3

Rime's latency documentation reports Mist v3 at 37 ms P50 and 56 ms P90 time to first audio with one concurrent request. At 12 concurrent requests, those published numbers remain 37 ms and 56 ms. The same page lets you compare Coda at 96 ms P50 and 98 ms P90 with one request, rising to 150 ms and 181 ms at 12.

That is more useful than a single best-case number because it exposes the tail and a load condition. It still is not your end-to-end call latency. Rime says typical network round trip adds 25 to 50 ms from much of the continental United States when the nearest region is used, while a far-coast route can add 60 to 85 ms. Your media server, audio buffer, and phone network come after that.

Mist is the speed choice; Coda is the broader, more metadata-rich option. Mist supports English, French, German, and Spanish with 94 voices. Coda supports eight production languages and 184 voices. Both stream over HTTP and WebSockets, but Coda offers word-level timestamps while Mist does not. If the application needs precise word boundaries to reconcile playback, that missing Mist metadata can outweigh its lower TTFA.

Best for: Teams that want percentile latency data and can operate inside four Mist languages.
Standout: Published P50 and P90 TTFA at both one and 12 concurrent requests.
Pricing: Mist v3 is $0.03 per 1,000 characters; Coda is $0.05 per 1,000.
Free trial: Yes, but Rime's live page conflicts on the quantity of included minutes.

Rime pricing, every current tier

Starter charges $0.03 per 1,000 characters for Mist and $0.05 for Coda, requires no credit card, and permits 20 concurrent TTS generations. Enterprise is custom and adds unlimited concurrent generations, unlimited custom voice clones, SLA and dedicated support, plus cloud, on-prem, or VPC deployment.

The page currently says “about 800 minutes free” near the Starter offer and “3,000 free minutes” in its FAQ. Treat the existence of signup credit as verified and its quantity as unsettled until the account shows a balance or Rime corrects the page. That inconsistency is small at production scale, but it is exactly why a pricing fact-lock is worth doing.

The upside
What it does well
4 points

  • Mist publishes both median and tail TTFA at stated concurrency.
  • The 37 ms P50 and 56 ms P90 figures are unusually fast.
  • HTTP and WebSocket streaming are available.
  • Enterprise supports cloud, VPC, or on-prem deployment.
The downside
Where it falls short
3 points

  • Mist has four production languages and no word-level timestamps.
  • Network placement can add more latency than the model itself.
  • The current pricing page conflicts on the amount of free signup usage.

Verdict: Rime Mist is the best speed specialist in the ranking. It moves above Deepgram only when the call flow can tolerate simpler interruption metadata and the real-region bakeoff confirms its percentile advantage.

3. Inworld Realtime TTS-2 Flash: Best for Global Cost and Language Coverage

Inworld Realtime TTS-2 Flash is the economic winner for a multilingual voice agent at public on-demand rates.

Inworld Realtime TTS-2 product page
Inworld Realtime TTS-2 Flash

Inworld's live TTS page publishes a 20 ms P90 time to first audio byte for Flash and 150 ms for the more expressive TTS-2. Those numbers are server-side and exclude network latency, so they cannot be placed directly beside Rime's total cloud guidance or Hume's typical first-audio figure. The useful conclusion is narrower: Inworld has built Flash for immediate streaming over WebSocket, and the model-side boundary is unlikely to be the largest part of a well-routed call.

Coverage is the stronger differentiator. Inworld lists 200+ languages and can clone a voice across those languages from 5 to 15 seconds of audio. That makes it attractive for a support agent that must keep one identity across markets without operating a separate vendor or voice per language.

The cost consequence is equally sharp. On-Demand Flash is $15 per million characters, compared with $30 per million for Rime Mist, $45 per million for post-promotion Deepgram Flux, and $50 per million for ElevenLabs Flash. At the 50-million-character scenario later in this article, that public rate is $750 per month before any volume discount.

The tradeoff is pricing complexity. Inworld uses monthly credits, model-specific usage rates, six plan levels, and different concurrency limits. A lower rate does not automatically mean a lower invoice when a paid tier's credit commitment exceeds actual usage. Match the tier to real consumption instead of buying a discount percentage.

Best for: High-volume multilingual agents and one voice identity across many languages.
Standout: 200+ languages, WebSocket streaming, and a $15-per-million-character on-demand Flash rate.
Pricing: Start free; Flash falls from $15 per million characters on On-Demand to below $5 on Enterprise.
Free trial: Yes. On-Demand includes up to 70 minutes of TTS.

Inworld pricing, every current tier

The live Inworld pricing page uses credits to cover TTS, STT, and language-model usage:

  • On-Demand: Starts free, includes up to 70 TTS minutes, charges Flash at $15 per million characters and TTS-2 at $25, with 5 concurrent requests.
  • Creator: $25 per month in credits, Flash at $10 per million and TTS-2 at $20, with 10 concurrent requests.
  • Builder: $100 per month in credits, Flash at $9 per million and TTS-2 at $17.50, with 50 concurrent requests.
  • Developer: $300 per month in credits, Flash at $8 per million and TTS-2 at $15, with 150 concurrent requests.
  • Growth: $1,500 per month in credits, Flash at $7 per million and TTS-2 at $12.50, with 500 concurrent requests.
  • Enterprise: Custom, with TTS-2 as low as $5 per million, Flash below $5 per million, and custom concurrency.

Inworld's own calculator assumes roughly 1,000 characters per generated minute. Use that convention for a first budget, then replace it with the measured character count from actual prompts. A terse collections agent and a verbose tutoring agent will not consume the same characters per minute.

The upside
What it does well
4 points

  • Flash publishes a 20 ms server-side P90 first-audio-byte figure.
  • Coverage spans 200+ languages with cross-lingual cloning.
  • Public character pricing is the lowest among the four cost finalists.
  • Concurrency scales from 5 requests on On-Demand to 500 on Growth.
The downside
Where it falls short
3 points

  • The latency figure excludes network transit.
  • Six credit tiers make the effective invoice harder to compare.
  • The expressive TTS-2 model costs more and publishes a higher latency boundary than Flash.

Verdict: Inworld is the first pick when the launch is global or character spend is a board-level line item. Deepgram stays ahead for English calls where native barge-in state saves more engineering and conversational risk than the price delta.

4. ElevenLabs Flash v2.5: Best Mature Multilingual Ecosystem

ElevenLabs Flash v2.5 is the safest mainstream choice when voice selection, multilingual operations, and a familiar API ecosystem matter together.

ElevenLabs Flash v2.5 model documentation
ElevenLabs Flash v2.5

ElevenLabs' model documentation publishes roughly 75 ms of latency for Flash v2.5, excluding application and network delay. The model supports 32 languages and a 40,000-character request limit. For a voice agent, the practical appeal is not one isolated number. It is the combination of low-latency positioning, broad language support, established voice operations, and a public pay-as-you-go path.

The important wall hides in text normalization. Flash disables number normalization by default to preserve latency. Phone numbers, dates, and currencies may be spoken in a way the caller does not expect. Enterprise customers can turn normalization on for v2.5. Everyone else should normalize those values in the language model before passing text into TTS.

That is not an edge case. A support agent regularly speaks order IDs, dates, amounts, addresses, confirmation codes, and phone numbers. The fastest voice is useless if a compact date string comes out as one giant number rather than a date. Put normalization into the prompt and test suite, not into a last-minute production patch.

ElevenLabs is also the one high-value partner in this roundup that fits the category honestly. It ranks fourth because the score follows the production decision, not the commercial relationship. It belongs above the remaining choices for a team that values a mature multilingual voice workflow, and below the top three when interruption state, transparent percentile reporting, or public character cost is the harder constraint.

Best for: Multilingual products that value an established voice ecosystem and simple API procurement.
Standout: Roughly 75 ms published model latency across 32 languages.
Pricing: Flash and Turbo cost $0.05 per 1,000 characters.
Free trial: Yes. Free / Pay As You Go includes 20,000 Flash or Turbo characters.

ElevenLabs pricing, every current tier

The live ElevenLabs API pricing page lists:

  • Free / Pay As You Go: $0 with 20,000 Flash or Turbo characters.
  • Starter: $6 per month with 120,000 characters.
  • Creator: $22 per month, with the first month at $11, and 440,000 characters.
  • Pro: $99 per month with 1,980,000 characters.
  • Scale: $299 per month with 5,980,000 characters.
  • Business: $990 per month with 19,800,000 characters.
  • Enterprise: Custom.

The Flash rate remains $0.05 per 1,000 characters across the displayed API tiers. That makes the plan decision about included usage, account features, and operational limits rather than a lower displayed per-character rate.

The upside
What it does well
4 points

  • Flash v2.5 supports 32 languages.
  • The published latency is about 75 ms before application and network delay.
  • The free tier gives a small but useful integration allowance.
  • A mature voice ecosystem reduces procurement and voice-management friction.
The downside
Where it falls short
3 points

  • Number, date, and currency normalization is off by default on Flash.
  • The $0.05-per-1,000-character rate is materially above Inworld Flash.
  • The headline latency excludes both application and network time.

Verdict: ElevenLabs is a defensible default, not the automatic number one. Pick it when voice operations and multilingual familiarity outweigh the extra character cost, then make normalization a release blocker.

5. Cartesia Sonic-3.6: Best for Persistent WebSocket Contexts

Cartesia Sonic-3.6 is the best fit for developers who want explicit WebSocket context control, cancellation, and timestamp events.

Cartesia Sonic-3.6 pricing page
Cartesia Sonic-3.6

Cartesia's live pricing table now names Sonic-3.6. Its WebSocket API carries a context identifier, and the same connection supports context cancellation plus word and phoneme timestamp responses. That is useful when an agent streams several clauses, needs to stop one utterance cleanly, and wants fine-grained playback metadata.

The honest latency finding is an omission: the current public Sonic-3.6 pages checked for this comparison do not publish a specific TTFA number. Older Cartesia latency figures attached to earlier Sonic versions should not be silently recycled onto 3.6. That keeps Cartesia below vendors with current, attributable latency boundaries, even though the transport design is compelling.

Pricing is credit-based and the displayed page was set to yearly billing. The minutes are approximate, which is fine for an initial budget but should be replaced by measured consumption before a contract.

Best for: Developers who need context IDs, cancellation, and timestamp-rich WebSocket output.
Standout: Word and phoneme timestamps plus explicit context cancellation.
Pricing: Free starts at about 27 minutes; paid annual-billing rates begin at $4 per month.
Free trial: Yes. The Free tier includes 20,000 monthly credits, shown as about 27 TTS minutes.

Cartesia pricing, every current tier

  • Free: $0 per month, about 27 TTS minutes, and 2 concurrent requests.
  • Pro: $4 per month on yearly billing, about 133 minutes, and 3 concurrent requests.
  • Startup: $39 per month on yearly billing, about 1,667 minutes, and 5 concurrent requests.
  • Scale: $239 per month on yearly billing, about 10,667 minutes, and 15 concurrent requests.
  • Enterprise: Custom credits, agent usage, volume pricing, and concurrency.

Cartesia's model naming and pricing are current enough to buy against. Its latency marketing is not specific enough to compare Sonic-3.6 confidently. Make the vendor earn that place with a real-region bakeoff.

The upside
What it does well
4 points

  • WebSocket contexts make cancellation and utterance state explicit.
  • Word and phoneme timestamps support precise playback logic.
  • The free plan is easy to prototype with.
  • Paid plans include unlimited workspace seats.
The downside
Where it falls short
3 points

  • No current public Sonic-3.6 TTFA number was found on the checked pages.
  • Displayed paid rates assume yearly billing.
  • Public concurrency tops out at 15 before Enterprise.

Verdict: Cartesia is a strong transport-first shortlist candidate. It rises only after its current model beats the leaders on your measured p95 path, because an old Sonic number is not a Sonic-3.6 fact.

6. Hume Octave 2: Best for Expressive Delivery

Hume Octave 2 is the expression specialist for a voice agent that must sound emotionally aware rather than merely fast.

Hume Octave 2 text-to-speech documentation
Hume Octave 2

Hume's Octave documentation marks Octave 2 as a preview. Hume publishes model latency around 100 ms excluding network transit, while instant mode typically makes first audio ready in roughly 200 ms depending on load and input complexity. Those are two different boundaries, and the second is closer to what an application feels before its own playback overhead.

The model supports English, Japanese, Korean, Spanish, French, Portuguese, Italian, German, Russian, Hindi, and Arabic during preview. It can clone a voice from as little as 15 seconds of audio. The API offers HTTP output streaming, bidirectional WebSocket streaming, and non-streaming responses.

Hume's reason to exist in this ranking is semantic and emotional delivery. A coaching agent, companion, or sensitive customer-care flow may benefit when emphasis and pacing carry meaning. A basic appointment reminder probably does not need to pay the same expression tax.

Best for: Coaching, companion, health-support, and character agents where delivery changes the experience.
Standout: Expressive synthesis with roughly 100 ms model latency in the Octave 2 preview.
Pricing: Free starts at about 10 minutes; paid plans run from $3 to $500 per month before Enterprise.
Free trial: Yes. Free includes 10,000 characters and 1 concurrent connection.

Hume pricing, every current tier

  • Free: $0 per month, 10,000 characters or about 10 minutes, 15 requests per minute, and 1 concurrent connection.
  • Starter: $3 per month, 30,000 characters or about 30 minutes, 15 requests per minute, and 5 connections.
  • Creator: $14 per month, first month $7, 140,000 characters or about 140 minutes, $0.15 per additional 1,000 characters, 75 requests per minute, and 5 connections.
  • Pro: $70 per month, 1 million characters or about 1,000 minutes, $0.12 per additional 1,000, 75 requests per minute, and 10 connections.
  • Scale: $200 per month, 3.3 million characters or about 3,300 minutes, $0.10 per additional 1,000, 150 requests per minute, and 20 connections.
  • Business: $500 per month, 10 million characters or about 10,000 minutes, $0.05 per additional 1,000, 225 requests per minute, and 30 connections.
  • Enterprise: Custom usage, rate limits, and concurrency.

The pricing curve is generous for exploration and expensive at Creator overage. At scale, Business eventually reaches $0.05 per 1,000 additional characters, the same displayed character rate as ElevenLabs Flash, but with a $500 monthly plan and a different voice proposition.

The upside
What it does well
4 points

  • Expressive delivery is the clearest differentiator in the field.
  • Both HTTP and bidirectional WebSocket streaming are available.
  • Voice cloning can begin with 15 seconds of audio.
  • The Free and Starter plans make an emotional-voice prototype inexpensive.
The downside
Where it falls short
3 points

  • Octave 2 remains in preview.
  • Typical first audio in instant mode is closer to 200 ms than the 100 ms model figure.
  • Creator overage is $0.15 per 1,000 characters.

Verdict: Hume is not the speed winner, and it does not need to be. Pick it when the agent's emotional delivery earns revenue, retention, or trust that a more neutral low-latency voice cannot.

7. Gradium: Best for a Unified EU and US Speech Stack

Gradium is the pragmatic choice when one provider for TTS and STT, automatic regional routing, and predictable bundles matter more than a published TTFA trophy.

Gradium voice AI homepage
Gradium

Gradium's API reference supports both REST and WebSocket speech endpoints. Requests go to one API hostname and are automatically routed to the nearest EU or US cluster. Gradium says EU traffic is processed in Europe and US traffic in the United States. That can simplify a regional architecture that would otherwise need separate endpoint selection.

The pricing model is monthly credits rather than a simple character rate. One TTS character consumes one credit, and Gradium maps about 45,000 characters to one hour of speech. The bundle can be attractive if the same account uses transcription or translation, but the comparison becomes less clean when TTS is the only workload.

The sharpest entry-tier wall is commercial rights. Free includes API access and about one hour of TTS, but commercial use is not allowed. The $13 XS plan is the first public commercial tier.

Best for: Products that want one regional vendor for synthesis and transcription.
Standout: One API hostname with automatic EU and US cluster routing.
Pricing: Free for noncommercial use; commercial plans start at $13 per month.
Free trial: Yes. Free includes 45,000 credits, about one TTS hour, and 2 TTS concurrency.

Gradium pricing, every current tier

  • Free: $0 per month, 45,000 credits, about 1 TTS hour, 2 concurrency, API access, and no commercial use.
  • XS: $13 per month, 225,000 credits, about 5 TTS hours, 5 concurrency, and commercial use.
  • S: $43 per month, 900,000 credits, about 20 TTS hours, 5 concurrency, and commercial use.
  • M: $340 per month, 9 million credits, about 200 TTS hours, 10 concurrency, and commercial use.
  • L: $1,615 per month, 45 million credits, about 1,000 TTS hours, 15 concurrency, and commercial use.
  • Enterprise: Custom with unlimited included credits and TTS plus custom concurrency.

Additional 100,000-credit packs cost $6.90 on XS, $5.00 on S, $4.00 on M, and $3.80 on L. That declining overage ladder rewards committing to a larger bundle, but it also means the effective rate depends on utilization.

The upside
What it does well
4 points

  • REST and WebSocket TTS live behind one API.
  • Traffic routes automatically to EU or US clusters.
  • TTS, STT, and translation share the same credit system.
  • The $13 XS plan gives commercial API access without a large commitment.
The downside
Where it falls short
3 points

  • The Free plan forbids commercial use.
  • The checked pages do not publish a precise current TTFA boundary.
  • Credit bundles make a TTS-only cost comparison less direct.

Verdict: Gradium is a stack-consolidation choice, not the latency benchmark winner. It makes the shortlist when regional data handling and using one speech vendor remove more operational work than a specialist saves in milliseconds.

8. OpenAI GPT-4o Mini TTS: Best When the Stack Is Already on OpenAI

OpenAI GPT-4o Mini TTS is the integration-convenience choice for a product already standardized on the OpenAI API.

OpenAI GPT-4o Mini TTS model documentation
OpenAI GPT-4o Mini TTS

The model accepts text and produces audio through the official speech endpoint. The API can stream audio or server-sent events, offers 13 named built-in voices plus custom voice IDs, and supports MP3, Opus, AAC, FLAC, WAV, and PCM output. Playback speed can range from 0.25 to 4.0.

The missing fact is a current official TTFA figure. OpenAI's model page does not publish one, so GPT-4o Mini TTS cannot rank as a low-latency leader on documentation alone. It belongs here because the endpoint can stream and because reusing an existing provider, key-management path, observability system, and procurement relationship has real value.

Pricing also resists a clean character comparison. OpenAI bills $0.60 per 1 million input text tokens and $12 per 1 million output audio tokens. That can be reasonable, but a buyer must forecast from observed token use rather than multiplying a public character rate. The model page lists a 2,000-input-token maximum.

Best for: Teams already using OpenAI that value one provider more than a documented latency lead.
Standout: Flexible audio formats, streaming modes, voice instructions, and provider consolidation.
Pricing: $0.60 per million input text tokens plus $12 per million output audio tokens.
Free trial: No. The model's Free usage tier is not supported.

OpenAI usage tiers, every current limit

The official GPT-4o Mini TTS model page lists no Free access. Tier 1 allows 500 requests per minute and 50,000 tokens per minute. Tier 2 allows 2,000 and 150,000. Tier 3 allows 5,000 and 600,000. Tier 4 allows 10,000 and 2 million. Tier 5 allows 10,000 and 8 million.

Those are usage limits, not monthly subscriptions. The production question is whether the token-metered speech bill and existing-stack simplicity beat a specialist's clearer character rate and stronger voice-agent protocol.

The upside
What it does well
4 points

  • The speech endpoint streams as audio or server-sent events.
  • Six output formats and 13 built-in voices cover many application paths.
  • Existing OpenAI infrastructure can reduce vendor and operational overhead.
  • Voice instructions provide delivery control without changing providers.
The downside
Where it falls short
3 points

  • No official current TTFA number is published.
  • Token-based pricing is harder to forecast beside character-based rivals.
  • Free usage is not supported.

Verdict: Keep OpenAI on the shortlist if it collapses a provider boundary. Do not call it the fastest until the real call path proves it.

What 50,000 Generated Minutes Cost

The cleanest budget comparison uses a common character assumption and public on-demand rates. Inworld's pricing page assumes about 1,000 characters per minute, so 50,000 generated minutes becomes 50 million characters. This is a scenario, not a promised invoice.

At that volume:

  • Inworld Flash On-Demand costs $750 per month at $15 per million characters.
  • Rime Mist v3 costs $1,500 per month at $0.03 per 1,000 characters.
  • Deepgram Flux Pay As You Go costs $2,250 per month after the promotion at $0.045 per 1,000 characters.
  • ElevenLabs Flash costs $2,500 per month at $0.05 per 1,000 characters.
Monthly TTS cost comparison at 50,000 generated minutes
At a common 50-million-character assumption, public rates span $750 to $2,500 per month.

Do not turn that gap into an automatic recommendation. A cheaper voice that increases repeat calls, misreads account numbers, or forces weeks of custom interruption work can cost more than it saves. Conversely, paying a familiar vendor premium without measuring a business benefit is just inertia.

The scenario has three deliberate simplifications. First, characters per minute vary with language, speech rate, and punctuation. Second, paid credit plans can create minimum spend or unlock lower usage rates. Third, enterprise discounts are private. Put one week of real agent transcripts through a character and token counter before signing an annual commitment.

Who Should Pick What

Pick Deepgram Flux TTS when English coverage is enough and callers regularly interrupt. Native spoken-state reporting is the decision-flipping capability.

Pick Rime Mist v3 when the fastest transparent percentile profile matters, the four supported languages cover the launch, and the application can work without Mist word timestamps. Put media servers near users, because the network can erase the model's advantage.

Pick Inworld TTS-2 Flash when the same agent must speak across many markets or when generated-character spend is the largest TTS constraint. Its public on-demand rate gives finance the most room in the 50,000-minute scenario.

Pick ElevenLabs Flash v2.5 when the organization wants a mature multilingual voice operation and a simple path from free usage to larger plans. Normalize every number, date, currency, code, and address before synthesis.

Pick Cartesia Sonic-3.6 when context cancellation and word or phoneme timing matter to the playback architecture. Require a fresh TTFA result from the real deployment, because the current model pages do not provide one to inherit.

Pick Hume Octave 2 when emotional delivery is part of the product rather than decoration. Accept preview risk and a higher first-audio boundary only when that expression changes the outcome.

Pick Gradium when one TTS and STT bill, automatic EU or US routing, and a regional speech stack remove operational work. Start on XS for any commercial prototype because Free is noncommercial.

Pick OpenAI GPT-4o Mini TTS when eliminating another vendor is worth more than having a specialist's latency documentation. Measure it; do not assume provider familiarity equals low end-to-end delay.

One question flips most close calls: What happens when the caller interrupts halfway through a consequential fact? If the answer is “the client guesses what played,” fix that before debating voice taste.

The Ones to Avoid for Latency-First Voice Agents

Pika Speech: Fast Full Clips, Wrong Live Transport Today

Pika Speech is an impressive narration and voice-cloning system, but its current API is the wrong transport for a latency-first live voice agent.

Pika Speech API documentation showing asynchronous job polling
Pika Speech

Pika's August 18, 2026 release reports a 0.02 real-time factor in locally run three-minute tests. Under its long-form benchmark, a minute of speech takes about 1.2 seconds to generate. The model produces 48 kHz audio, can clone from 5 seconds of reference audio, and accepts up to 5 minutes of speech per request.

That is excellent full-clip throughput. It is not evidence that a caller hears the first syllable quickly.

The live Pika Speech API documentation explicitly lists real-time streaming synthesis under “Not for.” A POST returns a queued job, and the client polls until the job is complete. The vendor's benchmark table records Pika Speech at $0.01 per generated minute as of August 14, 2026, making it attractive for bulk narration, but the asynchronous contract disqualifies it from this live-turn ranking today.

Best for: Narration, voiceover, and fast full-clip voice generation.
Standout: About 1.2 seconds to generate one minute under Pika's long-form benchmark.
Pricing: $0.01 per generated minute in the vendor's August benchmark table.
Free trial: Not confirmed on the checked model and release pages.

The upside
What it does well
3 points

  • Very fast long-form generation under the vendor benchmark.
  • 48 kHz output and short-reference voice cloning fit content production.
  • The listed $0.01-per-minute rate is unusually low.
The downside
Where it falls short
3 points

  • The current API is explicitly not for real-time streaming synthesis.
  • Generation is asynchronous and requires job polling.
  • Full-clip throughput does not establish time to first audio.

Verdict: Use Pika for queued speech jobs today. Re-evaluate it for live agents only when the public API exposes a real streaming path with a measured first-audio boundary.

Avoid “free unlimited” as a production requirement

Several providers offer real free credits or free monthly allowances. None of the checked pages promises unlimited, production-grade, low-latency TTS at no cost. Free plans also carry walls such as noncommercial use, low concurrency, unsupported access, or a small character balance.

Budget a paid path before the prototype works. Otherwise success itself becomes the incident.

Avoid the quality model when speed is the hard constraint

Most vendors now separate a fast conversational model from a more expressive or long-form model. Choose Flash over a slower ElevenLabs option for latency-first calls, Mist over Coda when raw TTFA is decisive, and Inworld Flash over TTS-2 when cost and speed matter more than maximum expression. Do not pay the latency tax for quality a phone line or use case cannot exploit.

The Monday Move

Next Monday, run a five-scenario bakeoff with Deepgram Flux, Rime Mist, Inworld Flash, and ElevenLabs Flash. Use the same language model output, media server, region, phone carrier, audio encoding, and playback buffer for every candidate.

The five scenarios should expose different failures:

  1. A short greeting that reveals first-audio delay.
  2. A long answer interrupted halfway through a sentence.
  3. A turn containing a person's name, an August date, a currency amount, phone number, and confirmation code.
  4. A streamed language-model response with uneven chunk sizes.
  5. A concurrency burst that reaches the expected launch peak.

Record text-ready time, first byte received, first frame queued, first frame played, p50 TTFA, p95 TTFA, interruption correctness, and the exact text the caller heard. Score pronunciation and voice preference only after every candidate clears the hard latency and barge-in threshold.

Then attach the real generated character count to the pricing tier the deployment would actually buy. The Monday decision is not “which demo sounds best?” It is “which survivor produces correct, interruptible calls at the lowest complete monthly cost?”

If Deepgram clears multilingual needs, its interruption state makes it the default. If the product is multilingual, start with Inworld and ElevenLabs. If raw first audio is the gating metric, keep Rime in the final pair. One day of disciplined measurement is more valuable than another week comparing vendor adjectives.

Frequently Asked Questions

What is the best TTS API?

Deepgram Flux TTS is the best overall API in this comparison for an English voice agent because it combines low published first audio with native interruption state. Inworld is the better answer when 200+ languages or the lowest public character rate matters more, while Rime is the speed specialist.

Is there a free low-latency text-to-speech API?

Yes. Deepgram offers a $200 Pay As You Go credit and a temporary Flux promotion through September 12, 2026. Inworld includes up to 70 TTS minutes on On-Demand, ElevenLabs includes 20,000 Flash characters, Cartesia includes about 27 minutes, Hume about 10 minutes, and Gradium about one hour for noncommercial use. These are prototype allowances, not free unlimited production plans.

What is the cheapest TTS API?

In the transparent 50-million-character scenario, Inworld Flash On-Demand is cheapest among the four character-priced finalists at $750 per month. Enterprise rates, credit minimums, language mix, and actual characters per minute can change the winner, so forecast from real transcripts before committing.

What is the best local TTS model?

A local or self-hosted model is a different procurement decision from a managed API. Deepgram Flux and Rime both document private deployment routes, but the winner depends on your hardware, concurrency, model access terms, and ability to operate the serving stack. Benchmark first audio and real-time factor on the exact hardware you intend to own.

Does a low-latency TTS API work on Android?

Yes. Keep the vendor API key on your backend, stream generated audio from that backend to the Android client, and measure first playback on the device. Embedding a secret in the app is unsafe, and mobile buffering or network conditions can dominate the provider's model-level latency.

Get the AI tools map for business owners to compare the rest of your agent stack without rebuilding this research from scratch.

Last Updated

Aug 26, 2026

CategoryAI
Newsletter

One letter, every Sunday. Working systems, not hot takes.

Build logs, working systems, and field notes from running a portfolio of AI ventures.

Weekly. No spam. Unsubscribe anytime.