Best Text to Speech in 2026: Speechify vs NaturalReader vs ElevenLabs vs Amazon Polly vs Google vs Azure vs OpenAI (Verified July 2026)

The best text-to-speech tools for reading, publishable narration, and APIs, with every price and limit verified in July 2026.

Wednesday, July 29, 2026Omid Saffari
Tools
  • SSpeechify
  • NNaturalReader
  • EElevenLabs
  • AAmazon Polly
  • GGoogle Cloud Text-to-Speech
  • AAzure AI Speech
  • OOpenAI GPT-4o mini TTS
Best Text to Speech in 2026: Speechify vs NaturalReader vs ElevenLabs vs Amazon Polly vs Google vs Azure vs OpenAI (Verified July 2026)

Speechify is the best text-to-speech app for listening, but its $29 Premium subscription does not buy commercial distribution rights. NaturalReader is the better document-first value at $79 a year, ElevenLabs is the $6 publishable-production pick, and Amazon Polly or Google Cloud wins when speech belongs inside software.

The first decision is not which voice sounds nicest in a demo. It is what the audio must become.

A reader app turns a PDF, webpage, email, or book into something you can listen to. A production studio turns a script into audio you can edit and publish. A speech API turns text into audio inside a product. Those jobs overlap at the synthesis engine, but their prices, rights, review steps, and failure modes are different.

Every plan, allowance, model, and limit below was checked against the vendors' live pages on 29 July 2026. This is a priced-and-analyzed comparison, not a claim that paid accounts were exercised or that the voices passed a controlled listening test.

The best text-to-speech tools at a glance

ToolBest forStarting priceFree trial
SpeechifyReading across devicesFree; Premium $29/monthFree plan
NaturalReaderDocuments, OCR, and educationFree; Lite $13.90/monthFree plan
ElevenLabsPublishable narrationFree; Starter $6/monthFree plan
Amazon PollyPredictable cloud speech$4 per 1M characters after free tierAWS free tier
Google Cloud Text-to-SpeechThe widest API price ladder$4 per 1M characters after free usageFree usage by model
Azure AI SpeechMicrosoft-stack volume commitments$15 per 1M characters on demandF0 gives 0.5M characters/month
OpenAI GPT-4o mini TTSPrompt-directed API speech$0.60 input and $12 audio output per 1M tokensNo API Free tier

Pick Speechify when the job is listening to whatever is already on your screen. Pick NaturalReader when PDFs, scanned pages, school deployment, and annual cost matter more than a polished cross-device experience. Pick ElevenLabs when the output will become a client video, course, audiobook, or campaign asset.

The choice flips to an API only when speech must be generated automatically inside software. Amazon Polly is the clean cost baseline. Google Cloud offers the broadest ladder from $4 Standard and WaveNet voices to $160 Studio voices. Azure earns consideration at committed enterprise volume. OpenAI is the direction-first choice when an instruction such as "calm, measured, with a restrained sense of urgency" is more useful than a conventional voice setting.

How these were picked

Seven products made the list because each owns a distinct buyer decision. The ranking turns on six criteria:

  1. Input workflow: Can it handle the material you already have, such as a webpage, PDF, scan, long script, or application response?
  2. Direction: Can you control pronunciation, pace, emotion, emphasis, and consistency, or only choose a voice and press play?
  3. Rights: Does the exact tier permit personal listening, commercial publishing, or product use?
  4. Cost unit: Is the meter words, characters, credits, text tokens, audio tokens, or a monthly allowance?
  5. Output workflow: Does the product include editing, export, review, and collaboration, or only return an audio file?
  6. The wall: What makes a sensible plan stop being sensible, such as a personal-use license, shared credit pool, per-model multiplier, minimum commitment, or missing price transparency?

The shortlist deliberately mixes reader apps, production studios, and APIs because buyers use the phrase "text to speech" for all three. It does not pretend they are interchangeable. Products that could only be described with voice-count marketing, an unverifiable price, or a vague quality adjective were cut.

Voice quality was not ranked. A credible listening comparison needs the same script, language, voice type, loudness, output format, delivery direction, and scoring rubric for every system. A homepage demo cannot establish long-form consistency, pronunciation accuracy, or whether a voice survives a demanding edit. The verdicts therefore rest on verifiable workflow, rights, controls, limits, and cost.

1. Speechify is the best text-to-speech app for listening

Speechify is the strongest first choice when your source material already exists and you want to hear it across the day rather than produce a distributable audio file.

Speechify text-to-speech Reader pricing page
Speechify

Its Reader product accepts the everyday shape of reading work: documents, cloud files, scanned material, and text on supported devices. Premium adds more than 1,000 voices across more than 60 languages, playback up to 5x, Scan & Listen, summaries, chats, and integrations with Google Drive, Dropbox, and Microsoft OneDrive. That combination makes it easier to move from a laptop article to a phone commute without rebuilding the material in a studio timeline.

The decisive limitation is rights, not voice quality. Speechify's live usage policy says Reader accounts cannot be used for commercial distribution without permission. The separate Speechify Studio product is the publishing path.

Best for: Cross-device listening to books, PDFs, articles, email, and cloud documents
Standout: The cleanest reading-first workflow in this group
Pricing: Reader Free $0; Reader Premium $29/month; Studio Free $0; Studio Starter $19/month; Studio Creator $49/month
Free trial: Reader and Studio both have ongoing free plans

The upside
What it does well
4 points

  • Reader Free establishes the workflow without a card
  • Premium supports more than 1,000 voices, more than 60 languages, and playback up to 5x
  • The 2026 Premium allowance reaches 1,000,000 Premium Voice words per month
  • Accessibility users can request extra usage for a documented need
The downside
Where it falls short
3 points

  • $29/month is expensive when document handling and annual value matter more than cross-device polish
  • Reader Premium does not permit commercial distribution without permission
  • Studio is a separate subscription, so upgrading Reader does not add publishing rights

Every Speechify Reader and Studio tier

The Reader pricing page publishes two tiers:

  • Reader Free: $0. Playback up to 1.5x, ten basic voices, listening across supported surfaces, and text-to-speech features only.
  • Reader Premium: $29/month. More than 1,000 voices, more than 60 languages, playback up to 5x, Scan & Listen, summaries, chats, cloud-drive integrations, voice typing, podcasts, and its Voice AI Assistant.

Speechify's usage policy adds the number that the price card leaves out. Premium subscribers are guaranteed 1,000,000 Premium Voice words each month through 31 December 2026. The contractual baseline below that temporary extension is 150,000 words per month. Usage resets monthly.

That is generous for personal listening. It still does not make Reader an audio-production license.

The Studio pricing page publishes three separate tiers:

  • Studio Free: $0 with 600 Studio credits, access to more than 1,000 voices, Voiceover Studio, Dubbing Studio, and Voice Changer. It does not include voice cloning or commercial usage rights.
  • Studio Starter: $19/month with 7,200 Studio credits, voice cloning, stock media, voiceover, dubbing, Voice Changer, and commercial usage rights.
  • Studio Creator: $49/month with 28,800 Studio credits and everything in Starter.

Studio meters different jobs at different rates. Voiceover consumes one credit per second. Dubbing consumes three. Avatar generation consumes 30. Re-exporting unchanged speech is free, but a new script or a change to pitch, speed, or emotional delivery triggers another generation and uses credits again.

A repeatable PDF-to-listening setup

The top pick should be easy to prove against your own material before paying. Use one representative file, not a polished demo.

  1. Choose the job before the file

    If the audio is only for you, open Reader. If another person will receive, stream, buy, or publish the audio, stop and open Studio or another commercial product instead.

  2. Use a difficult source

    Choose a PDF or webpage with headings, abbreviations, a table, and at least one scanned element. A clean paragraph proves almost nothing about a reading app.

  3. Set the listening pace

    Start at normal speed, then raise it only until comprehension begins to fall. Premium supports playback up to 5x, but the highest setting is a capability ceiling, not a sensible default.

  4. Check navigation, not just the voice

    Jump between headings, pause, resume on another device, and inspect how the app handles cloud files or scanned material. A reader earns its price by helping you move through the document, not by sounding impressive for one sentence.

  5. Audit the rights before export

    If the result now belongs in a public or paid asset, move the script into Studio. Reader Premium remains a personal listening product even after you pay $29/month.

When Speechify is worth it

Speechify is worth $29/month when reading follows you across devices, the 1,000,000-word 2026 allowance is useful, and friction costs more than the subscription. A lawyer moving between filings, a founder consuming reports during travel, or a reader who needs accessibility support can justify that convenience.

It is poor value when you mostly upload an occasional PDF on one computer. Twelve months at the published monthly rate costs $348 before any separate annual discount. NaturalReader Lite is $79/year, Plus is $119, and Pro is $159. The difference is large enough that the workflow needs to earn it.

2. NaturalReader is the best document-first value

NaturalReader is the better purchase when files, OCR, annotations, school deployment, and annual price matter more than the smoothest cross-device brand experience.

NaturalReader online document text-to-speech interface
NaturalReader

The personal app supports PDF, EPUB, DOCX, PPTX, Numbers, and Pages files. Paid tiers add OCR, which turns text inside an image or scan into machine-readable text, plus a Pronunciation Editor, annotations, AI Smart Filter, and MP3 conversion. Smart Filter matters because tables, page numbers, headers, and footers can turn an otherwise useful document into exhausting audio.

NaturalReader's biggest strength is also its biggest source of buyer error: Personal and Commercial are separate products. Every Personal tier, including paid MP3 export, is licensed for personal use only.

Best for: Document-heavy personal listening, OCR, and education deployments
Standout: The deepest document workflow at a lower annual price than Speechify Premium paid monthly
Pricing: Personal Free; Lite $13.90/month or $79/year; Commercial Starter $29/month or $198/year
Free trial: Free Personal plan plus 10,000 Commercial sample characters per day

The upside
What it does well
4 points

  • Free handles basic listening with a usable daily AI-voice allowance
  • Paid Personal tiers support OCR, common document formats, annotations, and MP3 conversion
  • Annual individual plans run from $79 to $159
  • EDU price bands and site licenses make procurement legible
The downside
Where it falls short
4 points

  • Personal and Commercial subscriptions do not include each other
  • Every Personal output is restricted to personal use
  • Commercial voice-provider multipliers can reduce the apparent character allowance by 10x or 20x
  • The product has enough tiers that buying without reading the rights page is risky

Every NaturalReader Personal tier

NaturalReader's Personal pricing page and voice-limit page publish the complete ladder:

  • Free: $0. Unlimited Free Voices, 20,000 Lite Voice characters per day, and 4,000 shared Plus and Cloned Voice characters per day. It cannot convert to MP3 and does not include OCR.
  • Lite: $13.90/month or $79/year. Unlimited Lite Voice listening, 1 million Lite Voice MP3 characters per month, OCR, common document formats, Pronunciation Editor, annotations, Smart Filter, and AI features. It does not include voice cloning, Plus Voices, or Pro Voices.
  • Plus: $20.90/month or $119/year. It adds 500,000 shared Plus and Cloned Voice listening characters per day, 1 million shared Plus and Cloned Voice MP3 characters per month, and up to two cloned voices.
  • Pro: $25.90/month or $159/year. It adds Pro Voices and Reading Styles while keeping the 500,000 shared listening-character daily limit and 1 million shared MP3-character monthly limit across Plus, Cloned, and Pro Voices.

The annual prices create a clean value ladder. Speechify Premium paid for 12 months at its published $29 monthly rate costs $348. NaturalReader Lite is $269 less, Plus is $229 less, and Pro is $189 less. Speechify can still win on workflow, but its convenience has a visible annual price.

Every NaturalReader education tier

NaturalReader sells two EDU group tiers with annual billing:

  • Premium EDU: $199 for one to five users, $299 for six to ten, $399 for 11 to 20, $499 for 21 to 30, $555 for 31 to 40, and $599 for 41 to 50. Above 50 users, pricing starts at $12 per user per year.
  • Plus EDU: $299 for one to five users, $499 for six to ten, $899 for 11 to 20, $1,200 for 21 to 30, $1,400 for 31 to 40, and $1,500 for 41 to 50. Above 50 users, pricing starts at $25 per user per year.
  • Site license: Custom price for a school with at least 2,000 enrolled users.

Pro is not available for EDU. Premium EDU excludes Plus and Pro Voices. Plus EDU adds Plus and Cloned Voice access but not Pro Voices.

This is the stronger institutional option when the requirement is reading support rather than content production. The rights boundary does not disappear in education: the Personal product remains for personal listening.

Every NaturalReader Commercial tier

NaturalReader Commercial is the separate path for audio that will be distributed. A free user can sample up to 10,000 characters each day. The current commercial plans are:

  • Starter: $29 per user monthly or $198 per user yearly, equivalent to $16.50/month billed annually. It includes 500,000 credits per month.
  • Creator: $49 per user monthly or $297 per user yearly, equivalent to $24.75/month billed annually. It includes 2 million credits per month.
  • Team: $33 per user monthly or $192 per user yearly, equivalent to $16/month billed annually. It requires at least two users and includes 2 million shared credits per user per month.

Every paid Commercial plan includes commercially licensed audio, access to Gemini, OpenAI, Azure, and ElevenLabs voice models, Prompt Control, more than 90 languages, up to four cloned voices, Script Assistant, Pronunciation Editor, and 44.1 kHz MP3 and WAV download.

The hidden cost is the credit multiplier. Gemini, OpenAI, Azure, Google Chirp HD, and legacy voices cost one credit per character. ElevenLabs Turbo costs ten. ElevenLabs HD costs 20.

That means Starter's 500,000 credits buy:

  • 500,000 characters with a one-credit voice
  • 50,000 characters with ElevenLabs Turbo
  • 25,000 characters with ElevenLabs HD

Creator and Team provide 2 million credits, which become 2 million, 200,000, or 100,000 characters under the same three multipliers.

Who should skip NaturalReader

Skip NaturalReader Personal if you need commercial output. Skip NaturalReader Commercial if you want one predictable character meter across every model. Skip the whole product family if a clean mobile handoff matters more than document controls and the extra annual cost of Speechify does not concern you.

NaturalReader wins when the source document is the difficult part. It is less convincing when the brief starts as a clean script and the output needs performance direction.

3. ElevenLabs is the best text-to-speech tool for publishable narration

ElevenLabs is the strongest production pick when the audio must be directed, exported, and published rather than merely listened to.

ElevenLabs text-to-speech product page
ElevenLabs

Starter costs $6/month and is the first tier here that combines a commercial license with Instant Voice Cloning and Dubbing Studio. Creator adds Professional Voice Cloning. Pro adds 44.1 kHz PCM through the API and 192 kbps audio. The platform covers the space between a browser studio and an API without forcing a buyer to assemble both on day one.

That breadth creates the main wall. Text to speech, speech to text, music, sound effects, voice changing, isolation, and dubbing all draw from one shared credit pool. A plan that appears to include 121 narration minutes can deliver less when the same account also runs dubbing or repeated generations.

Best for: Narration, audiobooks, courses, branded video, dubbing, and directed voice assets
Standout: Commercial production starts at $6/month
Pricing: Free $0; Starter $6; Creator $22; Pro $99; Scale $299; Business $990; Enterprise custom
Free trial: Free plan with 10,000 credits and about ten TTS minutes

The upside
What it does well
4 points

  • Starter adds a commercial license and Instant Voice Cloning at $6/month
  • The plan ladder reaches browser creators, teams, API users, and enterprises
  • Paid unused credits can roll over for up to two months
  • Pro and above publish clear higher-quality output options
The downside
Where it falls short
4 points

  • Every product competes for one shared credit pool
  • Free does not list the commercial license included on Starter
  • Direction changes can require another generation and more credits
  • A studio workflow is unnecessary overhead for someone who only wants a PDF read aloud

Every ElevenLabs tier

The live pricing page lists:

  • Free: $0/month, 10,000 credits, about ten TTS minutes, and three Studio projects.
  • Starter: $6/month, 30,000 credits, about 30 TTS minutes, commercial license, Instant Voice Cloning, 20 Studio projects, commercial music use, and Dubbing Studio.
  • Creator: $22/month, with the first month at $11, 121,000 credits, about 121 TTS minutes, Professional Voice Cloning, and additional credits.
  • Pro: $99/month, 600,000 credits, about 600 TTS minutes, 44.1 kHz PCM API output, and 192 kbps audio.
  • Scale: $299/month, 1.8 million credits, about 1,800 TTS minutes, three workspace seats, team collaboration, and three Professional Voice Clones.
  • Business: $990/month, 6 million credits, about 6,000 TTS minutes, ten seats, ten Professional Voice Clones, and low-latency TTS listed as low as $0.05/minute.
  • Enterprise: Custom pricing, credits, and seats, plus custom DPA and SLA terms, HIPAA BAAs, custom SSO, elevated concurrency, managed dubbing, volume discounts, and priority support.

Annual billing charges the equivalent of ten months. That puts Starter at $5/month, Creator at $18.33, Pro at $82.50, Scale at $249.17, and Business at $825.

Paid unused credits roll over for up to two months and up to twice the monthly quota. With the current month's allotment, the balance can reach three times the monthly quota. Free credits do not roll over.

The simple subscription cost per included UI minute settles near one band after Starter. Starter works out to $0.20 per included minute, Creator about $0.18, and Pro, Scale, and Business about $0.17. That is before another ElevenLabs product consumes the same pool.

The same brief before and after speech editing

A written sentence and a spoken sentence do different work. Written copy can carry parentheses, dense lists, slashes, URLs, and visual hierarchy because the reader can stop and scan. Audio disappears as it is heard.

Consider a fictional product brief:

Written copy: The Studio collection includes the Atlas chair, Arc lamp, and Field table, available in bone, cobalt, and moss.

The speech-ready version creates audible structure:

Speech-ready copy: Meet the Studio collection. First, the Atlas chair. Then, the Arc lamp. Finally, the Field table. Choose bone, cobalt, or moss.

The second version is not smarter prose. It has shorter semantic units, explicit sequence words, and fewer chances for the listener to lose the list. A pronunciation dictionary should then lock any brand, product, person, or technical term whose spelling does not reliably reveal its sound.

This is the craft bar that a polished demo hides. A voice can sound convincing and still make a long script difficult to follow.

A ship-ready narration recipe

  1. Shape the script for ears. Break visual lists into audible sequence. Replace raw URLs, unexplained abbreviations, and parenthetical clauses.
  2. Lock pronunciation. Add brand names, people, acronyms, and domain-specific words to the pronunciation workflow before generating a long passage.
  3. Generate a short representative block. Include a list, an emotional transition, and one difficult name. Do not spend the full allowance proving an easy opening sentence.
  4. Review against the script. Listen for omissions, mispronunciation, unintended emphasis, breathless pacing, and a tone that changes the meaning.
  5. Export only after rights are clear. Starter is the first ElevenLabs tier that lists the commercial license. Source-voice consent remains a separate requirement from the subscription.

For a deeper comparison of production studios such as Murf, Hume, Speechify Studio, WellSaid, Respeecher, Cartesia, and Inworld, use the AI voice generator buyer's guide.

4. Amazon Polly is the best predictable speech API

Amazon Polly is the cleanest cost baseline when a product needs speech on demand and the team already knows how to operate an AWS service.

Amazon Polly text-to-speech cloud service page
Amazon Polly

The service charges by character and lets customers cache and replay generated speech without an additional Polly charge. Speech Marks can return timing metadata for uses such as synchronized highlighting or animation. That makes Polly more than a file generator, but it remains infrastructure: the reader, timeline, editorial review, and distribution workflow are yours to build.

Best for: Applications that need predictable per-character billing and AWS operations
Standout: A transparent $4 to $100 per million-character ladder
Pricing: Standard $4, Neural $16, Generative $30, Long-Form $100 per 1M characters
Free trial: AWS free tier varies by voice class

The upside
What it does well
4 points

  • Straight per-character pricing is easy to budget
  • Generated speech can be cached and replayed without another Polly charge
  • Speech Marks support synchronized product experiences
  • Four voice classes let cost and output requirements be separated
The downside
Where it falls short
4 points

  • It is not a reader app or production studio
  • The team must build ingestion, editing, review, storage, and playback around it
  • Free allowances for Neural, Generative, and Long-Form voices apply for the first 12 months
  • A lower unit price does not establish better voice quality

Every Amazon Polly voice class and free allowance

The Polly pricing page publishes four paid classes:

  • Standard: $4 per 1 million characters.
  • Neural: $16 per 1 million characters.
  • Generative: $30 per 1 million characters.
  • Long-Form: $100 per 1 million characters.

The free tier includes:

  • 5 million Standard characters per month
  • 1 million Neural characters per month for the first 12 months
  • 100,000 Generative characters per month for the first 12 months
  • 500,000 Long-Form characters per month for the first 12 months

AWS uses 1 million characters as approximately 23 hours and eight minutes of speech in its own examples. At that scale, the four classes cost $4, $16, $30, or $100 before storage, delivery, application code, monitoring, and human review.

When Polly wins

Polly wins when cost predictability and AWS integration matter more than natural-language direction. A product that reads notifications, articles, lessons, or status updates can budget character volume before it picks a voice class.

It loses when an editor needs to direct a performance in plain language, collaborate on a timeline, or make many subjective retakes. Those jobs turn a cheap API into an internal tool-building project.

5. Google Cloud Text-to-Speech has the widest API price ladder

Google Cloud Text-to-Speech is the best infrastructure choice when you want one vendor to offer low-cost legacy voices, current Chirp voices, instant custom voice, premium Studio voices, and prompt-directed Gemini-TTS.

Google Cloud Text-to-Speech product page
Google Cloud Text-to-Speech

No other ranked API spans so many billing models. Character-priced voices run from $4 to $160 per million characters. Gemini-TTS uses separate text-input and audio-output token prices. That range is powerful, but it also makes "Google TTS costs X" an incomplete statement.

Best for: Teams that want several speech model classes inside Google Cloud
Standout: The broadest ladder from $4 Standard to $160 Studio and token-priced Gemini-TTS
Pricing: $4 to $160 per 1M characters, or token-priced Gemini-TTS
Free trial: Free usage varies by model; Gemini-TTS and Instant Custom Voice list none

The upside
What it does well
4 points

  • Several model classes cover basic, neural, generative, custom, and studio needs
  • Standard and WaveNet include 4 million free characters
  • Chirp 3 HD gives a current $30 per million-character option
  • Gemini-TTS accepts text-based direction
The downside
Where it falls short
4 points

  • The price model changes between characters and tokens
  • Spaces, line breaks, and most SSML tags count toward character billing
  • Studio is 40 times the $4 Standard rate after free usage
  • The API still leaves the reader, editor, review, and rights workflow to the buyer

Every current Google text-to-speech price

The live Google pricing page separates Gemini-TTS, current TTS models, and legacy models.

Gemini-TTS

  • Gemini 2.5 Flash TTS and Gemini 2.5 Flash-Lite Preview TTS: No free allowance. $0.50 per 1 million text input tokens plus $10 per 1 million audio output tokens.
  • Gemini 3.1 Flash TTS Preview: No free allowance. $1 per 1 million text input tokens plus $20 per 1 million audio output tokens.
  • Gemini 2.5 Pro TTS: No free allowance. $1 per 1 million text input tokens plus $20 per 1 million audio output tokens.

Google defines an audio token as one twenty-fifth of a second for these prices. Output duration therefore matters directly.

Current character-priced models

  • Chirp 3 HD: 1 million free characters, then $30 per 1 million.
  • Instant Custom Voice: No free allowance, then $60 per 1 million.

Legacy character-priced models

  • WaveNet: 4 million free characters, then $4 per 1 million.
  • Standard: 4 million free characters, then $4 per 1 million.
  • Neural2: 1 million free characters, then $16 per 1 million.
  • Polyglot Preview: 1 million free characters, then $16 per 1 million.
  • Studio: 1 million free characters, then $160 per 1 million.

The $160 Studio rate is not a rounding error. It is 40 times the $4 Standard rate. A team should hear a material benefit on its own script before routing production volume into that tier.

Where the choice flips

Pick Standard or WaveNet when cost and free volume dominate. Pick Neural2 when the $16 tier is acceptable and a current neural class fits the language and voice requirement. Pick Chirp 3 HD when the $30 class earns its premium. Pick Instant Custom Voice only when a custom identity is worth $60 per million characters before the surrounding consent and review work.

Gemini-TTS is a different purchase. It is useful when prompt-based delivery control matters, but token billing makes it harder to compare directly with a character meter. Budget the output duration, not just the script length.

6. Azure AI Speech is best for Microsoft-stack volume commitments

Azure AI Speech earns its place when the organization already buys through Azure and can use a large character commitment, not because its on-demand price wins.

Azure AI Speech product page
Azure AI Speech

Standard on-demand Neural and Neural HD Flash speech costs $15 per million characters. That is 3.75 times the $4 Standard rate from Amazon Polly or Google Cloud. The gap narrows through commitments, but the first published commitment starts at 80 million characters.

Best for: Azure organizations with procurement leverage and predictable high volume
Standout: Commitments lower the effective rate from $15 to as little as $7.50 per 1M characters
Pricing: F0 free; Standard $15 per 1M characters; commitments from $960
Free trial: F0 includes 0.5M Neural characters per month

The upside
What it does well
4 points

  • F0 provides a recurring 0.5 million-character Neural allowance
  • Standard supports on-demand real-time and batch synthesis
  • Commitment pricing becomes competitive at large volume
  • Custom Voice pricing publishes synthesis, training, and hosting separately
The downside
Where it falls short
4 points

  • $15 per million characters loses to $4 Standard APIs at low volume
  • The first commitment requires 80 million characters and $960
  • Custom Voice hosting can dominate synthesis cost
  • The pricing surface includes many speech products, making procurement easy to misread

Every core Azure TTS tier

The Azure AI Speech pricing page publishes:

  • Free F0: 0.5 million Neural Text to Speech characters per month.
  • Standard S0: Neural and Neural HD Flash cost $15 per 1 million characters for real-time or batch synthesis.
  • Professional Custom Voice: Standard synthesis costs $24 per 1 million characters. Neural HD synthesis costs $48 per 1 million. Training costs $52 per compute hour, capped at $936 per training. Endpoint hosting costs $4.04 per model per hour.
  • 80 million-character commitment: $960, equivalent to $12 per 1 million.
  • 400 million-character commitment: $3,900, equivalent to $9.75 per 1 million.
  • 2 billion-character commitment: $15,000, equivalent to $7.50 per 1 million.

The on-demand comparison is simple. Azure Standard at $15 is close to AWS Neural and Google Neural2 at $16, not to their $4 Standard class.

The custom-voice comparison is different. A Professional Custom Voice endpoint hosted continuously for a 30-day, 720-hour month costs $2,908.80 before it synthesizes one character. That can make sense for an approved, heavily used brand voice. It is wasteful for occasional narration.

When Azure wins

Azure wins when speech belongs inside an existing Microsoft architecture, procurement already has an Azure agreement, and volume is predictable enough to use a commitment. It also provides a legible cost model for approved Professional Custom Voice deployments.

It loses for a small product that wants the cheapest on-demand baseline or a creator who needs an editor instead of cloud infrastructure.

7. OpenAI GPT-4o mini TTS is best for prompt-directed API speech

OpenAI GPT-4o mini TTS is the most direct API choice when delivery should follow natural-language instruction and the team is comfortable with token billing.

OpenAI GPT-4o mini TTS model documentation
OpenAI GPT-4o mini TTS

The model accepts direction for accent, emotional range, intonation, impressions, speaking speed, tone, and whispering. The Audio API can return MP3, Opus, AAC, FLAC, WAV, or PCM, with WAV or PCM recommended for the fastest response. OpenAI requires a clear disclosure that the voice is AI-generated rather than human.

The wall is cost comparability. GPT-4o mini TTS charges separately for text input tokens and audio output tokens. A slower or longer delivery changes output cost even when the script does not change. Converting that price to "per million characters" would require assumptions the character-priced APIs do not need.

Best for: Prompt-directed speech inside an application
Standout: Delivery can be shaped with natural-language instructions
Pricing: $0.60 per 1M text input tokens plus $12 per 1M audio output tokens
Free trial: The API Free tier is not supported

The upside
What it does well
4 points

  • Natural-language instructions control several performance dimensions
  • Six common output formats cover streaming, compressed delivery, and editing
  • The current model page publishes token prices and a 2,000-input-token limit
  • Custom voices require a consent recording and a separate sample
The downside
Where it falls short
4 points

  • Token pricing is harder to compare with character-priced APIs
  • Free API usage is not supported
  • The 2,000-input-token limit requires long scripts to be segmented
  • The live guide contradicts itself on the built-in voice count

Current price, limits, and voice options

The GPT-4o mini TTS model page lists:

  • Text input: $0.60 per 1 million tokens
  • Audio output: $12 per 1 million tokens
  • Maximum input: 2,000 tokens
  • Free usage tier: Not supported

The text-to-speech guide calls GPT-4o mini TTS its newest and most reliable text-to-speech model. It recommends marin or cedar for best quality.

The same live page contains a documentation mismatch worth noticing. Its introduction says the Audio API comes with 11 built-in voices. The detailed list contains 13 names: alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin, and cedar. Use the enumerated list when selecting a voice and treat the earlier count as stale copy.

Custom voices are limited to eligible customers. Creating one requires a consent recording plus a separate sample recording, and an organization can create at most 20 custom voices.

An exact prompt-directed speech recipe

  1. Choose the current model

    Use gpt-4o-mini-tts. Keep each request within the documented 2,000-input-token maximum.

  2. Choose the voice before the direction

    Start with marin or cedar, the two voices OpenAI recommends for best quality. Hold the voice constant while changing the delivery instruction.

  3. Write one delivery instruction

    Use a compact instruction such as: "Speak with calm authority, moderate pace, restrained emotion, and a short pause after each section heading." This is a direction, not evidence of a generated result.

  4. Use an edit-friendly output

    Choose WAV when the audio will enter an editing timeline, or PCM when a low-latency application can handle raw samples. MP3 is the default for general delivery.

  5. Disclose and review

    Tell listeners the voice is AI-generated, then review the audio against the approved script for omissions, pronunciation, emphasis, and meaning.

The model is ship-ready as an engine, not as a complete editorial system. Your application still needs script management, retry logic, storage, playback, review, and disclosure.

What one million characters costs

Character-priced APIs are easy to normalize because the unit stays constant:

  • $4: Amazon Polly Standard, Google Standard, and Google WaveNet
  • $7.50: Azure at a 2 billion-character commitment
  • $9.75: Azure at a 400 million-character commitment
  • $12: Azure at an 80 million-character commitment
  • $15: Azure Standard on demand
  • $16: Amazon Polly Neural, Google Neural2, and Google Polyglot Preview
  • $24: Azure Professional Custom Voice synthesis
  • $30: Amazon Polly Generative and Google Chirp 3 HD
  • $48: Azure Professional Custom Voice Neural HD synthesis
  • $60: Google Instant Custom Voice
  • $100: Amazon Polly Long-Form
  • $160: Google Studio
Physical column chart comparing four text-to-speech API prices per million characters
The same million characters can cost $4, $16, $30, or $160 before workflow costs

The $4 tier is the cost baseline, not an automatic quality winner. The $160 Google Studio tier costs 40 times as much. A buyer needs a controlled listening result on the same script to justify that spread.

OpenAI and Gemini-TTS sit outside this ladder because they bill text input and audio output tokens. Output duration and delivery change the audio meter. A per-character conversion without a fixed tokenization and speaking-duration assumption would create false precision.

Reader and studio economics are also different:

  • Speechify Premium costs $348 for 12 months at the published $29 monthly rate, before any separate annual discount.
  • NaturalReader costs $79/year for Lite, $119 for Plus, or $159 for Pro.
  • ElevenLabs included-minute cost moves from $0.20 on Starter to about $0.18 on Creator and about $0.17 on Pro, Scale, and Business, before other products consume the shared credits.
  • NaturalReader Commercial's Starter allowance can shrink from 500,000 characters to 25,000 when the selected voice costs 20 credits per character.

Who should pick what

The reliable decision rule is listen, publish, or build.

Pick Speechify to listen. It wins when material follows you across devices and the $29/month convenience is worth more than NaturalReader's annual savings. The choice flips to NaturalReader when files, OCR, annotations, EDU deployment, or cost dominate.

Pick NaturalReader for difficult documents. Free is useful, Lite adds OCR and MP3 conversion, and Plus or Pro increases voice access. The choice flips away the moment the audio must be distributed. Then use NaturalReader Commercial or a production studio.

Pick ElevenLabs to publish. Starter creates the cleanest low-cost path to a commercial license, Instant Voice Cloning, and a studio. The choice flips to Speechify Studio when you already prefer its creator workflow, or to a cloud API when speech is generated automatically at product scale.

Pick Amazon Polly to build cheaply and predictably. It is the baseline for an AWS team. The choice flips to Google when you need its wider model ladder, to Azure when Microsoft procurement and high-volume commitments matter, or to OpenAI when natural-language performance direction is the defining feature.

Pick Google Cloud for model range. Standard and WaveNet cover the $4 baseline, Neural2 the $16 band, Chirp the $30 band, and Studio the $160 band. The choice flips when that range creates more procurement and evaluation work than value.

Pick Azure for committed enterprise volume. On-demand Standard at $15 is not the low-cost choice. The case improves at 80 million, 400 million, or 2 billion committed characters.

Pick OpenAI for prompt-directed application speech. It wins when the delivery instruction is a product feature. The choice flips to a character-priced API when predictable unit economics matter more than flexible delivery.

Three-path decision flow routing text-to-speech buyers to read, publish, or build
Choose the job first: read, publish, or build

The ones to avoid

The products and plans below are not universally bad. They are bad purchases under a specific condition.

Avoid Speechify Reader and NaturalReader Personal for commercial output. Both products can produce audio. Neither paid reader subscription turns personal listening into a commercial distribution license.

Avoid free creator tiers for client work unless commercial rights are explicit. Speechify Studio Free has no commercial usage rights. ElevenLabs adds its commercial license on Starter. A free workflow audition is not permission to ship.

Hold PlayHT when price transparency is part of the decision. Its official pricing URL returned no usable live public table through the required fact lock. A commercial comparison should not fill that gap with cached third-party prices.

Send Murf buyers to the creative-studio category. Murf's current page publishes Free, Creator at $19/month billed $228 annually, Business at $66/month billed $792 annually, and Enterprise. It is better compared with other production studios in the AI voice generator guide than with document readers.

Avoid choosing by voice count. Speechify publishes more than 1,000 voices, OpenAI lists 13 names, and cloud providers offer model families rather than a single library number. None of those counts establishes pronunciation, long-form consistency, or fit for your script.

Avoid custom voice deployment without a permission record. OpenAI requires a consent recording and a separate sample for custom voices. The same operational standard should govern every provider: record who owns the source, what the clone may say, where it may run, and when permission ends.

Frequently asked questions

What is the best free text-to-speech tool?

NaturalReader is the best free document reader in this group. Free includes unlimited basic voices, 20,000 Lite Voice characters per day, and 4,000 shared Plus and Cloned Voice characters per day, but no MP3 conversion or OCR. Speechify Free is the easier cross-device audition, with ten basic voices and playback up to 1.5x.

What is the best text-to-speech app?

Speechify is the best text-to-speech app for most personal listening because it combines device coverage, cloud integrations, Scan & Listen, and playback up to 5x. NaturalReader is the better value when OCR, file handling, annotations, and a $79 to $159 annual plan matter more.

What is the best text-to-speech tool for students?

NaturalReader is the strongest student and school option. Its paid Personal tiers handle common document formats, OCR, annotations, pronunciation, and Smart Filter, while Premium EDU and Plus EDU publish price bands from small groups to more than 50 users. The site-license path starts at 2,000 enrolled users.

Can text-to-speech audio be used commercially?

Yes, when the exact tier grants commercial use and the underlying script and source voice are cleared. ElevenLabs Starter, Speechify Studio Starter, and NaturalReader Commercial publish commercial paths. Speechify Reader and NaturalReader Personal do not.

What does OpenAI text-to-speech cost?

GPT-4o mini TTS costs $0.60 per 1 million text input tokens plus $12 per 1 million audio output tokens. The API Free tier is unsupported, and each request accepts up to 2,000 input tokens.

How much does Amazon Polly text to speech cost?

Amazon Polly costs $4 per 1 million characters for Standard voices, $16 for Neural, $30 for Generative, and $100 for Long-Form. AWS says 1 million characters is approximately 23 hours and eight minutes of speech in its examples.

How much does Google Cloud Text-to-Speech cost?

Google starts at $4 per 1 million characters for Standard and WaveNet after free usage. Neural2 and Polyglot cost $16, Chirp 3 HD costs $30, Instant Custom Voice costs $60, and Studio costs $160. Gemini-TTS uses separate text-input and audio-output token prices.

How much does Azure text to speech cost?

Azure F0 includes 0.5 million Neural characters per month. Standard on-demand Neural and Neural HD Flash cost $15 per 1 million characters. Commitments lower the effective rate to $12, $9.75, or $7.50 per million at 80 million, 400 million, or 2 billion committed characters.

Is text to speech the same as speech to text?

No. Text to speech turns written input into audio. Speech to text turns spoken audio into written output. A voice assistant may use both, but they solve opposite conversion problems and are billed separately.

Can I convert text to speech online for free?

Yes. Speechify and NaturalReader both offer free browser access, while ElevenLabs offers a free creator tier. Free limits and rights differ: NaturalReader Free cannot export MP3, Speechify Reader is not for commercial distribution without permission, and ElevenLabs lists its commercial license on Starter.

Get the AI business workflow audit checklist to document the source text, usage rights, pricing tier, voice, disclosure, review owner, and final approval before generated speech enters a live campaign or product.

Last Updated

Jul 29, 2026

CategoryDesign

More from Design

View all Design articles
Newsletter

One letter, every Sunday. Working systems, not hot takes.

Build logs, working systems, and field notes from running a portfolio of AI ventures.

Weekly. No spam. Unsubscribe anytime.