Text-to-Speech AI Models Report 2026: 16 Models Compared

Global Text-to-Speech AI Models Research Report 2026

Global Text-to-Speech AI Models Research Report

16 Leading Models Compared — September 2026

Quality · Latency · Pricing · Scenario Recommendations

Executive Summary

The 2026 TTS market is no longer a one-dimensional race for “who sounds most human.” It has split into five fronts — quality, latency, cost, controllability, and compliance. ElevenLabs and Inworld lead on overseas quality, Cartesia and ByteDance’s Doubao Voice dominate low-latency real-time use, Google and OpenAI compete on price-performance, while China’s open-source camp (Qwen3-TTS, CosyVoice 3, IndexTTS-2) delivers Apache 2.0 licensing + sub-100ms latency + free self-hosting — the best value available to developers worldwide. There is no all-round champion. Picking by scenario is the only correct approach.

Key Findings:

  • Overall quality ceiling: Inworld TTS-1.5 Max (~1236 ELO)
  • Latency leaders: ElevenLabs Flash v2.5 (75ms), Cartesia Sonic 3 (90ms)
  • Best for Chinese: Volcano Engine Doubao Voice (ranked #1 in SuperCLUE, July 2026)
  • Best value: Google Gemini Flash (~$6/M tokens), OpenAI ($15-30/M chars)
  • Unique for video dubbing: IndexTTS-2 (millisecond duration control)
  • Top open-source for commercial use: Qwen3-TTS, CosyVoice 3 (Apache 2.0)
  • Go-global poster child: MiniMax Speech 2.8 HD (global top five)

Model Overview (Commercial APIs)

Model Vendor Latest Version Open Source API Price
ElevenLabs ElevenLabs Eleven v3 (2026) No Yes $120-220 / M chars
Inworld Inworld AI Realtime TTS-2 / TTS-1.5 Max No Yes Usage-based (mid-high)
Cartesia Cartesia Sonic 3.5 No Yes ~$50 / M chars
OpenAI TTS OpenAI gpt-4o-mini-tts (2026) No Yes $15-30 / M chars
Google Google Gemini 3.1 Flash TTS No Yes $6-30 / M chars
Volcano Engine Doubao ByteDance Seed-TTS 2.0 (2026) No Yes ~¥1.3 / 1K chars
Alibaba Qwen-Audio Alibaba Cloud Qwen-Audio-3.0-TTS Partial Yes Bailian tiered
iFLYTEK Super-Human iFLYTEK Super-Human Synthesis (2026) No Yes Enterprise quote
MiniMax MiniMax Speech 2.8 No Yes $60-100 / M chars

Model Overview (Open Source)

Model Org License Size Time-to-First-Audio
Qwen3-TTS Alibaba Apache 2.0 0.6B / 1.7B 97ms (end-to-end)
CosyVoice 3 Alibaba Apache 2.0 0.5B ~150ms first chunk
IndexTTS-2 Bilibili Apache 2.0 Undisclosed Offline synthesis
Fish Speech FishAudio Non-commercial — Under 150ms
F5-TTS Community MIT 330M Streaming
GPT-SoVITS Community MIT — Non-streaming (3-5s)
Kokoro-82M hexgrad Apache 2.0 82M Under 0.3s
Higgs Audio V2 BosonAI Apache 2.0 3B —
Moshi Kyutai Apache 2.0 7B Full-duplex real-time

Detailed Model Analysis

Overseas Commercial APIs

ElevenLabs Eleven v3 Quality Benchmark

Vendor: ElevenLabs
Latest: Eleven v3 (2026)
Latency: 75ms (Flash v2.5)
Price: $120-220 / M chars

Best-in-class voice cloning and voice-library ecosystem — 30 seconds of reference audio is enough for instant cloning across 32 languages. In an independent naturalness test, pronunciation accuracy hit 81.97% (vs. OpenAI’s 77.30%) and prosody accuracy 64.57% (OpenAI 45.83%); Flash v2.5 at ~75ms brings it into the real-time tier. The quality ceiling — and the most expensive tier.

Pros

  • #1 in voice cloning and voice library
  • 32 languages, leading accuracy
  • Flash v2.5 at 75ms enables real-time
  • Ideal for building brand voice assets

Cons

  • 8-11× the price of OpenAI
  • 300-600ms latency on long text
  • Facing a BIPA class action (May 2026)

Inworld Realtime TTS-2 #1 on Realtime

Vendor: Inworld AI
Latest: Realtime TTS-2 / TTS-1.5 Max
Latency: P90 <250ms (Mini <130ms)
Price: Usage-based (mid-high)

Ranks #1 on the Artificial Analysis realtime leaderboard; TTS-1.5 Max scores ~1236 ELO. Zero-shot cloning (5-15s audio) at no extra charge; Realtime TTS-2 supports 8-dimensional natural-language style control (emotion, pitch, volume, pace, etc.) and H100/B200 on-prem deployment, covering 100+ languages (15 GA).

Pros

  • #1 on realtime, very low P90 latency
  • Zero-shot cloning at no extra cost
  • 8-dimensional natural-language control
  • H100/B200 on-prem deployment

Cons

  • Only 15 production-grade languages
  • Long-tail languages weaker than ElevenLabs’ 32
  • Brand and tooling ecosystem still young

Cartesia Sonic 3.5 Latency King

Vendor: Cartesia
Latest: Sonic 3.5 (2026)
Latency: ~90ms (Turbo 40ms)
Price: ~$50 / M chars

With a TTFA of ~90ms, it sets the standard for real-time agents. Its SSM (state-space model) architecture makes inference cost scale linearly with context (quadratically for Transformers), giving it a cost-structure edge at scale. 3-second instant cloning, 40+ languages, at roughly one-third of ElevenLabs’ price.

Pros

  • ~90ms latency, real-time benchmark
  • SSM architecture scales cost-effectively
  • Roughly one-third of ElevenLabs’ price
  • 3-second instant cloning

Cons

  • ~1054 ELO, a tier behind leaders
  • Emotional depth traded for speed
  • Access from mainland China is unfriendly

OpenAI gpt-4o-mini-tts Best Value

Vendor: OpenAI
Latest: gpt-4o-mini-tts (2026)
Latency: 200-400ms
Price: $15-30 / M chars

Extreme value with natural-language style control (“sound more excited”). Same SDK as the GPT ecosystem; the Realtime API supports full-duplex dialogue and paralinguistics (laughter, hesitation) across 57+ languages. Best for teams already inside the OpenAI ecosystem.

Pros

  • $15-30 / M chars, very low cost
  • Natural-language style control
  • Same SDK as the GPT ecosystem
  • Realtime API supports full-duplex

Cons

  • Only 13 built-in voices, no cloning
  • No SSML, 4096-char input cap
  • Lower pronunciation/prosody than ElevenLabs
  • ~1106 ELO — “good enough,” not a benchmark

Google Gemini 3.1 Flash TTS Cheapest at Top Tier

Vendor: Google
Latest: Gemini 3.1 Flash TTS / Chirp 3
Latency: Moderate
Price: $6-30 / M chars

At ~$6/M tokens, the Flash tier is the cheapest among top-quality engines; Gemini 3.1 Flash TTS sits in the Arena’s top tier. New GCP customers get $300 in credits plus 1M free characters per month, with enterprise SLAs and global nodes — ideal for large-scale narration and global deployments.

Pros

  • Cheapest among top-quality tiers
  • Gemini 3.1 Flash TTS in Arena’s top tier
  • Global nodes + enterprise SLA
  • 1M free characters per month

Cons

  • Fragmented lineup (Chirp 3 HD ~$30)
  • Weaker style control/cloning than specialists
  • Chinese dialects and polyphones lag domestic vendors

Other Overseas Players Worth Watching

Hume Octave: the emotion specialist (it can “act”), suited to emotion-first companion apps, though mid-table on quality overall.

Amazon Polly / Azure TTS: the safe choice for enterprise legacy ecosystems. Azure Neural is ~$15-16/M chars with 500K free chars/month and measured first-chunk latency of ~120ms in China; Polly suits high-throughput AWS workloads, but naturalness now trails the new generation by a tier.

xAI Grok Voice (launched July 2026): a speech-to-speech bundle (telephony, cloning, 80+ voices) at $0.05/min, attacking Cartesia/ElevenLabs’ agent market on price; but it is a voice-agent model rather than a dubbing engine, and its maturity is unproven.

Speechify SIMBA 3.0: broke into the global Arena top ten at very low cost — a dark horse for accessibility/listening scenarios.

Chinese Commercial APIs

Volcano Engine Doubao Voice (Seed-TTS 2.0) #1 in Chinese

Vendor: ByteDance
Latest: Seed-TTS 2.0 (2026)
Latency: <300ms (WebSocket streaming)
Price: ~¥1.3 / 1K chars

Ranked #1 overall in SuperCLUE’s Chinese speech leaderboard (70.81, July 2026). First-chunk latency under 300ms, with a measured streaming frame-interval standard deviation of 57ms (vs. iFLYTEK’s 113ms) — roughly 53% lower risk of dropouts. Sharing the Doubao LLM ecosystem, it enables end-to-end speech dialogue chains with instruction-based emotion control and paralinguistics.

Pros

  • #1 overall in SuperCLUE Chinese speech
  • <300ms first chunk, stable streaming
  • ~¥1.3/1K chars + full multilingual SDK
  • Same ecosystem as Doubao LLM

Cons

  • Legacy vs. LLM voices differ widely — test both
  • Deeply tied to ByteDance’s ecosystem
  • English and other languages weaker than its Chinese

Alibaba Cloud Qwen-Audio-3.0-TTS Open Source + Managed

Vendor: Alibaba Cloud (Bailian)
Latest: Qwen-Audio-3.0-TTS (2026)
Latency: ~300ms (Flash)
Price: Bailian tiered

Second overall in SuperCLUE (68.28). Supports 16 languages + 20 Chinese dialects and sentence-level emotion tags like [gasp] and [angry]. Its companion open-source Qwen3-TTS (Apache 2.0, 97ms end-to-end, 3-second zero-shot cloning) lets teams “validate before paying” — ideal for high-volume, cost-sensitive teams.

Pros

  • Second overall in SuperCLUE
  • 16 languages + 20 dialects, rich emotion tags
  • Open-source version is Apache 2.0
  • Validate before paying

Cons

  • Many product lines raise selection/migration cost
  • Managed vs. self-hosted quality differs
  • No unified spec — benchmark it yourself

iFLYTEK Super-Human Synthesis Government & Enterprise

Vendor: iFLYTEK
Latest: Super-Human Synthesis (2026)
Latency: Official P50 180ms (measured >1500ms)
Price: Enterprise quote

Chinese MOS of 4.5+/5.0, close to human; the most complete dialect coverage in the industry (20+ dialects + 8 foreign languages with real-time switching). Strong channels in education, government, healthcare, and automotive, with full government/enterprise compliance credentials and MRCP support for legacy banking/government systems.

Pros

  • Chinese MOS close to human
  • Most complete dialect coverage
  • Compliance credentials + MRCP legacy support
  • Deep education/government/healthcare channels

Cons

  • Closed API, requires enterprise vetting
  • Opaque pricing, poor SMB onboarding
  • Third-party first-chunk >1500ms — far from the official claim
  • Throughput only ~60% of Volcano Engine

MiniMax Speech 2.8 Global Top Five

Vendor: MiniMax
Latest: Speech 2.8 HD / Turbo
Latency: 400ms+
Price: $60-100 / M chars

China’s go-global poster child — the HD version’s 1164 ELO puts it in the global top five. 40+ languages with rich paralinguistics (laughter, breathing, sighs); supports pinyin/IPA pronunciation coverage, word-level timestamps, and emotion parameters, fitting production workflows well. Turbo at $60/M chars offers better value than ElevenLabs.

Pros

  • HD version in the global top five
  • 40+ languages + rich paralinguistics
  • Word-level timestamps + emotion parameters
  • Turbo beats ElevenLabs on value

Cons

  • HD at ~$100/M chars is still pricey
  • 400ms+ latency rules out real-time agents
  • Domestic vs. overseas pricing differs

Other Chinese Commercial Players

Tencent Cloud TTS: long-text API supports 100K characters + SSML + async tasks, tied to the WeChat ecosystem, Video Accounts, and Tencent Zhiying digital-human pipelines; its weakness is that its advanced models are less aggressive than Volcano/Alibaba.

Baidu AI Cloud: voice cloning supports Chinese, English, and Japanese plus some dialects and emotion parameters, winning via Baidu Cloud bundling; its weakness is that dialects and emotion parameters cannot be set together on some endpoints.

StepFun Step-Audio 2.5: a unified audio-language foundation model (ASR/TTS/Realtime in one) that drops the encoder-adapter for a pure LLM backbone with generative-reward RLHF; 67.6% Arena win rate. Its weakness is a less mature commercial API ecosystem than Volcano/Alibaba.

Shared positioning: Tencent Cloud and Baidu Cloud are “ecosystem-locked” players — choosing them is often choosing an entire cloud, not just TTS.

Open-Source Models (China-Led)

IndexTTS-2 Unique for Video Dubbing

Org: Bilibili
License: Apache 2.0 (commercial OK)
VRAM: ~12GB
Languages: Chinese/English + a few

Decouples emotion from timbre and offers millisecond-level duration control — a capability unique to it for lip-sync dubbing — plus Qwen3-tuned natural-language emotion instructions and WER as low as 1.6. For video/short-drama dubbing with second-precise timing, it is the only model that natively delivers.

Pros

  • Millisecond duration control (unique)
  • Emotion–timbre decoupling
  • Qwen3 natural-language emotion instructions
  • Apache 2.0, commercial-friendly

Cons

  • ~12GB VRAM
  • Only Chinese/English + a few languages
  • Modest generation speed

Qwen3-TTS Best Open-Source for Commercial

Org: Alibaba
License: Apache 2.0 (commercial OK)
Size: 0.6B / 1.7B
Latency: 97ms (end-to-end)

97ms end-to-end latency, 3-second zero-shot cloning, and instruction-level control under a free-for-commercial Apache 2.0 license. With sub-100ms latency and consumer-GPU feasibility, it is one of the lowest-barrier options for self-hosting a real-time voice agent.

Pros

  • 97ms end-to-end latency
  • 3-second zero-shot cloning
  • Apache 2.0, free for commercial use
  • Instruction-level control

Cons

  • Narrower language coverage than commercial
  • Modest high-fidelity ceiling

CosyVoice 3 Top Chinese Open-Source

Org: Alibaba
License: Apache 2.0 (commercial OK)
Size: 0.5B
Latency: ~150ms first chunk

9 languages + 18 Chinese dialects, ~150ms streaming first chunk, runnable on 8GB VRAM, with Chinese CER among the industry’s lowest. The go-to open-source choice for Chinese self-hosting, forming a “free-for-commercial tier” together with Qwen3-TTS.

Pros

  • 9 languages + 18 Chinese dialects
  • ~150ms first chunk, streaming-ready
  • Runs on 8GB VRAM
  • Chinese CER among the lowest

Cons

  • High-fidelity ceiling below diffusion-based models
  • Fewer languages than top commercial engines

Other Open-Source Models at a Glance

Fish Speech: strong cloning quality, sub-150ms latency, active community — but its open weights cannot be used commercially; you must go through its API.

F5-TTS: MIT-licensed, 330M and lightweight, good flow-matching naturalness, low deployment barrier; cloning similarity is average (SS 0.779).

GPT-SoVITS: MIT-licensed, best few-shot fine-tuning clone fidelity, full VITS toolchain; non-streaming (3-5s), so offline dubbing only.

Kokoro-82M: Apache 2.0, the speed champion (<0.3s), runs on CPU; few languages, no cloning, mid-tier quality.

Higgs Audio V2: Apache 2.0, trained on ~10M hours, strong multi-speaker dialogue expressiveness; resource-hungry and research-oriented.

Moshi: Apache 2.0, the open-source full-duplex benchmark (dual-stream modeling + Inner Monologue); a dialogue model rather than a dubbing engine, mid-tier audio quality.

ChatTTS / VibeVoice: ChatTTS has good conversational naturalness but restricted licensing; VibeVoice supports 90-minute long-form and 4 speakers but was disabled in August 2025 over misuse risk — research use only.

Comparison Matrix: The Whole Field at a Glance

Platform Arena ELO Time-to-First-Audio Price (per M chars) Languages Cloning
Inworld TTS-1.5 Max ~1236 P90 <250ms Mid-high 100+ (15 GA) 5-15s free
ElevenLabs Eleven v3 1178 75ms (Flash) $120-220 32 30s instant
MiniMax Speech 2.8 HD 1164 400ms+ $60-100 40+ Zero-shot
OpenAI TTS ~1106 200-400ms $15-30 57+ None
Cartesia Sonic 3.5 ~1054 90ms ~$50 40+ 3s
Google Chirp 3 Top tier Moderate $6-30 40+ Limited
Volcano Engine Doubao SuperCLUE #1 (Chinese) <300ms ~¥1.3/1K chars (≈$13) Chinese-first Yes
Alibaba Qwen-Audio-3.0-TTS SuperCLUE #2 ~300ms Bailian tiered 16 + 20 dialects 3s
iFLYTEK Super-Human Chinese MOS 4.5+ Official P50 180ms (measured >1500ms) Enterprise quote 20+ dialects Yes

Note: ELO figures are a snapshot of the Artificial Analysis blind-test leaderboard (H1 2026); rankings shift weekly, so treat them as ranges, not fixed values. iFLYTEK latency reflects third-party measurements that diverge from official claims — always benchmark before buying.

Scenario Recommendations

Chinese Real-Time Dialogue / Outbound Agent

Volcano Engine Doubao Voice

#1 in Chinese on SuperCLUE + <300ms first chunk + ~¥1.3/1K chars. The overall default for Chinese developers.

English / Multilingual Real-Time Agent

Cartesia Sonic 3.5

TTFA of 90ms with an SSM architecture that scales cost-effectively. Alternatives: Inworld, ElevenLabs Flash.

Audiobooks / High-Quality Dubbing

ElevenLabs Eleven v3

#1 in cloning quality and voice library; the quality premium is recoverable through content monetization. Alternative: MiniMax Speech 2.8 HD.

Multilingual Content Going Global

MiniMax Speech 2.8

Rich paralinguistics + 40+ languages + word-level timestamps; fits production workflows well.

Video / Short-Drama Dubbing (Lip-Sync, Timing)

IndexTTS-2

The only Apache 2.0 model with millisecond duration control — the only one that natively handles lip-sync and second-precise timing.

Education / Government / Finance (Channel & Compliance)

iFLYTEK

Most dialects + government/enterprise credentials + MRCP legacy compatibility. Alternatives: Tencent Cloud / Baidu Cloud.

Data-Sensitive / Cost-Sensitive Self-Hosting

Qwen3-TTS + CosyVoice 3

Apache 2.0, free for commercial use, runnable on 8GB VRAM. Lightweight alternative: F5-TTS.

Near-Zero-Budget Prototyping

Kokoro-82M

Runs on CPU, speaks in 0.3 seconds. The cheapest tier beyond OpenAI’s free credits.

Decision Framework

1. Chinese real-time dialogue? → Volcano Engine Doubao Voice

2. English / multilingual real-time agent? → Cartesia Sonic 3.5

3. Highest possible quality? → ElevenLabs / Inworld

4. Multilingual content going global? → MiniMax Speech 2.8

5. Video dubbing with timing / lip-sync? → IndexTTS-2

6. Education / government / finance (compliance-heavy)? → iFLYTEK

7. Self-hosting to control cost? → Qwen3-TTS / CosyVoice 3

8. Rock-bottom cost? → OpenAI / Google Flash

Technology Trends

1. “Low-frame-rate tokenizer + LLM semantic tokens + Flow Matching acoustics” is becoming the mainstream paradigm. Baichuan-Audio, Qwen-Audio-3.0, and CosyVoice 2 have all converged on this route — discrete tokens for understanding, continuous representations for high-fidelity synthesis.

2. Full-duplex real-time dialogue is the hottest battleground. Moshi’s dual-stream modeling is the open-source reference architecture; GPT-4o Realtime, Doubao Realtime, and Gemini Live are pushing end-to-end latency into the 300ms range on the closed side. Turn-taking and overlapping-speech modeling are the core challenges.

3. Emotion–timbre decoupling is becoming standard. IndexTTS-2 uses a gradient-reversal layer to separate “who is speaking” from “what emotion,” paired with natural-language emotion instructions; SSML is being replaced by natural-language control.

4. Audio watermarking and embedded compliance. Directions such as in-codec watermarking reflect regulation moving upstream — speech generation is shifting from a “capability problem” to a “traceability problem.”

Compliance & Legal Risk: Two Tracks, China and the U.S.

China: Under the Interim Measures for the Management of Generative AI Services and the Provisions on the Administration of Deep Synthesis of Internet Information Services, TTS-generated content requires visible/invisible labeling, public-facing services must complete filing, and voice cloning requires authorization from the person being cloned. When using domestic commercial APIs, platform-side compliance is clear; when self-hosting open-source models, filing and labeling obligations fall on the user.

United States: In May 2026, leaders such as ElevenLabs faced a BIPA class action in Illinois (alleging training data was used without voice-owner consent); by spring 2026, 46 U.S. states had enacted deepfake-related laws. When procuring overseas TTS services, assess vendors’ training-data compliance statements and indemnification terms.

Key Takeaways

1. Five fronts, no all-round champion. Quality, latency, cost, controllability, and compliance each form their own front line. Choose by real scenario, not by who tops the leaderboard.

2. The quality gap has largely closed. The blind-test top five span the U.S. (Inworld, ElevenLabs, Google) and China (MiniMax) — the leaders are now within the noise range of one another.

3. Latency is a hard metric. The full-chain budget is ~450ms (ASR 100 + LLM first token 150 + TTS first audio 120 + buffer 40); a TTS component over 500ms is unusable in real-time dialogue.

4. Open source has risen — China leads. The Qwen3-TTS / CosyVoice 3 / IndexTTS-2 trio “Apache 2.0 + sub-100ms + consumer GPU” is now the world’s strongest tier.

5. Licenses are a hidden trap. Fish Speech and ChatTTS open weights are not commercially usable, and VibeVoice has been disabled — such traps are more fatal than quality gaps. Verify the license chain before signing.

6. Compliance is a long-term variable. China’s deep-synthesis labeling obligations and the U.S. BIPA litigation plus 46-state deepfake laws will reshape the boundaries of cloning businesses.

© 2026 DeepForgeHub Research. Data sourced from Artificial Analysis, SuperCLUE, IndexTTS technical reports, and public benchmarks and hands-on testing. Model versions and pricing as of September 2026.

Newsletter Updates

Enter your email address below and subscribe to our newsletter

Leave a Reply

Your email address will not be published. Required fields are marked *