Global Text-to-Speech AI Models Research Report
16 Leading Models Compared — September 2026
Quality · Latency · Pricing · Scenario Recommendations
Executive Summary
The 2026 TTS market is no longer a one-dimensional race for “who sounds most human.” It has split into five fronts — quality, latency, cost, controllability, and compliance. ElevenLabs and Inworld lead on overseas quality, Cartesia and ByteDance’s Doubao Voice dominate low-latency real-time use, Google and OpenAI compete on price-performance, while China’s open-source camp (Qwen3-TTS, CosyVoice 3, IndexTTS-2) delivers Apache 2.0 licensing + sub-100ms latency + free self-hosting — the best value available to developers worldwide. There is no all-round champion. Picking by scenario is the only correct approach.
Key Findings:
- Overall quality ceiling: Inworld TTS-1.5 Max (~1236 ELO)
- Latency leaders: ElevenLabs Flash v2.5 (75ms), Cartesia Sonic 3 (90ms)
- Best for Chinese: Volcano Engine Doubao Voice (ranked #1 in SuperCLUE, July 2026)
- Best value: Google Gemini Flash (~$6/M tokens), OpenAI ($15-30/M chars)
- Unique for video dubbing: IndexTTS-2 (millisecond duration control)
- Top open-source for commercial use: Qwen3-TTS, CosyVoice 3 (Apache 2.0)
- Go-global poster child: MiniMax Speech 2.8 HD (global top five)
Model Overview (Commercial APIs)
| Model | Vendor | Latest Version | Open Source | API | Price |
|---|---|---|---|---|---|
| ElevenLabs | ElevenLabs | Eleven v3 (2026) | No | Yes | $120-220 / M chars |
| Inworld | Inworld AI | Realtime TTS-2 / TTS-1.5 Max | No | Yes | Usage-based (mid-high) |
| Cartesia | Cartesia | Sonic 3.5 | No | Yes | ~$50 / M chars |
| OpenAI TTS | OpenAI | gpt-4o-mini-tts (2026) | No | Yes | $15-30 / M chars |
| Gemini 3.1 Flash TTS | No | Yes | $6-30 / M chars | ||
| Volcano Engine Doubao | ByteDance | Seed-TTS 2.0 (2026) | No | Yes | ~¥1.3 / 1K chars |
| Alibaba Qwen-Audio | Alibaba Cloud | Qwen-Audio-3.0-TTS | Partial | Yes | Bailian tiered |
| iFLYTEK Super-Human | iFLYTEK | Super-Human Synthesis (2026) | No | Yes | Enterprise quote |
| MiniMax | MiniMax | Speech 2.8 | No | Yes | $60-100 / M chars |
Model Overview (Open Source)
| Model | Org | License | Size | Time-to-First-Audio |
|---|---|---|---|---|
| Qwen3-TTS | Alibaba | Apache 2.0 | 0.6B / 1.7B | 97ms (end-to-end) |
| CosyVoice 3 | Alibaba | Apache 2.0 | 0.5B | ~150ms first chunk |
| IndexTTS-2 | Bilibili | Apache 2.0 | Undisclosed | Offline synthesis |
| Fish Speech | FishAudio | Non-commercial | — | Under 150ms |
| F5-TTS | Community | MIT | 330M | Streaming |
| GPT-SoVITS | Community | MIT | — | Non-streaming (3-5s) |
| Kokoro-82M | hexgrad | Apache 2.0 | 82M | Under 0.3s |
| Higgs Audio V2 | BosonAI | Apache 2.0 | 3B | — |
| Moshi | Kyutai | Apache 2.0 | 7B | Full-duplex real-time |
Detailed Model Analysis
Overseas Commercial APIs
ElevenLabs Eleven v3 Quality Benchmark
Best-in-class voice cloning and voice-library ecosystem — 30 seconds of reference audio is enough for instant cloning across 32 languages. In an independent naturalness test, pronunciation accuracy hit 81.97% (vs. OpenAI’s 77.30%) and prosody accuracy 64.57% (OpenAI 45.83%); Flash v2.5 at ~75ms brings it into the real-time tier. The quality ceiling — and the most expensive tier.
Pros
- #1 in voice cloning and voice library
- 32 languages, leading accuracy
- Flash v2.5 at 75ms enables real-time
- Ideal for building brand voice assets
Cons
- 8-11× the price of OpenAI
- 300-600ms latency on long text
- Facing a BIPA class action (May 2026)
Inworld Realtime TTS-2 #1 on Realtime
Ranks #1 on the Artificial Analysis realtime leaderboard; TTS-1.5 Max scores ~1236 ELO. Zero-shot cloning (5-15s audio) at no extra charge; Realtime TTS-2 supports 8-dimensional natural-language style control (emotion, pitch, volume, pace, etc.) and H100/B200 on-prem deployment, covering 100+ languages (15 GA).
Pros
- #1 on realtime, very low P90 latency
- Zero-shot cloning at no extra cost
- 8-dimensional natural-language control
- H100/B200 on-prem deployment
Cons
- Only 15 production-grade languages
- Long-tail languages weaker than ElevenLabs’ 32
- Brand and tooling ecosystem still young
Cartesia Sonic 3.5 Latency King
With a TTFA of ~90ms, it sets the standard for real-time agents. Its SSM (state-space model) architecture makes inference cost scale linearly with context (quadratically for Transformers), giving it a cost-structure edge at scale. 3-second instant cloning, 40+ languages, at roughly one-third of ElevenLabs’ price.
Pros
- ~90ms latency, real-time benchmark
- SSM architecture scales cost-effectively
- Roughly one-third of ElevenLabs’ price
- 3-second instant cloning
Cons
- ~1054 ELO, a tier behind leaders
- Emotional depth traded for speed
- Access from mainland China is unfriendly
OpenAI gpt-4o-mini-tts Best Value
Extreme value with natural-language style control (“sound more excited”). Same SDK as the GPT ecosystem; the Realtime API supports full-duplex dialogue and paralinguistics (laughter, hesitation) across 57+ languages. Best for teams already inside the OpenAI ecosystem.
Pros
- $15-30 / M chars, very low cost
- Natural-language style control
- Same SDK as the GPT ecosystem
- Realtime API supports full-duplex
Cons
- Only 13 built-in voices, no cloning
- No SSML, 4096-char input cap
- Lower pronunciation/prosody than ElevenLabs
- ~1106 ELO — “good enough,” not a benchmark
Google Gemini 3.1 Flash TTS Cheapest at Top Tier
At ~$6/M tokens, the Flash tier is the cheapest among top-quality engines; Gemini 3.1 Flash TTS sits in the Arena’s top tier. New GCP customers get $300 in credits plus 1M free characters per month, with enterprise SLAs and global nodes — ideal for large-scale narration and global deployments.
Pros
- Cheapest among top-quality tiers
- Gemini 3.1 Flash TTS in Arena’s top tier
- Global nodes + enterprise SLA
- 1M free characters per month
Cons
- Fragmented lineup (Chirp 3 HD ~$30)
- Weaker style control/cloning than specialists
- Chinese dialects and polyphones lag domestic vendors
Other Overseas Players Worth Watching
Hume Octave: the emotion specialist (it can “act”), suited to emotion-first companion apps, though mid-table on quality overall.
Amazon Polly / Azure TTS: the safe choice for enterprise legacy ecosystems. Azure Neural is ~$15-16/M chars with 500K free chars/month and measured first-chunk latency of ~120ms in China; Polly suits high-throughput AWS workloads, but naturalness now trails the new generation by a tier.
xAI Grok Voice (launched July 2026): a speech-to-speech bundle (telephony, cloning, 80+ voices) at $0.05/min, attacking Cartesia/ElevenLabs’ agent market on price; but it is a voice-agent model rather than a dubbing engine, and its maturity is unproven.
Speechify SIMBA 3.0: broke into the global Arena top ten at very low cost — a dark horse for accessibility/listening scenarios.
Chinese Commercial APIs
Volcano Engine Doubao Voice (Seed-TTS 2.0) #1 in Chinese
Ranked #1 overall in SuperCLUE’s Chinese speech leaderboard (70.81, July 2026). First-chunk latency under 300ms, with a measured streaming frame-interval standard deviation of 57ms (vs. iFLYTEK’s 113ms) — roughly 53% lower risk of dropouts. Sharing the Doubao LLM ecosystem, it enables end-to-end speech dialogue chains with instruction-based emotion control and paralinguistics.
Pros
- #1 overall in SuperCLUE Chinese speech
- <300ms first chunk, stable streaming
- ~¥1.3/1K chars + full multilingual SDK
- Same ecosystem as Doubao LLM
Cons
- Legacy vs. LLM voices differ widely — test both
- Deeply tied to ByteDance’s ecosystem
- English and other languages weaker than its Chinese
Alibaba Cloud Qwen-Audio-3.0-TTS Open Source + Managed
Second overall in SuperCLUE (68.28). Supports 16 languages + 20 Chinese dialects and sentence-level emotion tags like [gasp] and [angry]. Its companion open-source Qwen3-TTS (Apache 2.0, 97ms end-to-end, 3-second zero-shot cloning) lets teams “validate before paying” — ideal for high-volume, cost-sensitive teams.
Pros
- Second overall in SuperCLUE
- 16 languages + 20 dialects, rich emotion tags
- Open-source version is Apache 2.0
- Validate before paying
Cons
- Many product lines raise selection/migration cost
- Managed vs. self-hosted quality differs
- No unified spec — benchmark it yourself
iFLYTEK Super-Human Synthesis Government & Enterprise
Chinese MOS of 4.5+/5.0, close to human; the most complete dialect coverage in the industry (20+ dialects + 8 foreign languages with real-time switching). Strong channels in education, government, healthcare, and automotive, with full government/enterprise compliance credentials and MRCP support for legacy banking/government systems.
Pros
- Chinese MOS close to human
- Most complete dialect coverage
- Compliance credentials + MRCP legacy support
- Deep education/government/healthcare channels
Cons
- Closed API, requires enterprise vetting
- Opaque pricing, poor SMB onboarding
- Third-party first-chunk >1500ms — far from the official claim
- Throughput only ~60% of Volcano Engine
MiniMax Speech 2.8 Global Top Five
China’s go-global poster child — the HD version’s 1164 ELO puts it in the global top five. 40+ languages with rich paralinguistics (laughter, breathing, sighs); supports pinyin/IPA pronunciation coverage, word-level timestamps, and emotion parameters, fitting production workflows well. Turbo at $60/M chars offers better value than ElevenLabs.
Pros
- HD version in the global top five
- 40+ languages + rich paralinguistics
- Word-level timestamps + emotion parameters
- Turbo beats ElevenLabs on value
Cons
- HD at ~$100/M chars is still pricey
- 400ms+ latency rules out real-time agents
- Domestic vs. overseas pricing differs
Other Chinese Commercial Players
Tencent Cloud TTS: long-text API supports 100K characters + SSML + async tasks, tied to the WeChat ecosystem, Video Accounts, and Tencent Zhiying digital-human pipelines; its weakness is that its advanced models are less aggressive than Volcano/Alibaba.
Baidu AI Cloud: voice cloning supports Chinese, English, and Japanese plus some dialects and emotion parameters, winning via Baidu Cloud bundling; its weakness is that dialects and emotion parameters cannot be set together on some endpoints.
StepFun Step-Audio 2.5: a unified audio-language foundation model (ASR/TTS/Realtime in one) that drops the encoder-adapter for a pure LLM backbone with generative-reward RLHF; 67.6% Arena win rate. Its weakness is a less mature commercial API ecosystem than Volcano/Alibaba.
Shared positioning: Tencent Cloud and Baidu Cloud are “ecosystem-locked” players — choosing them is often choosing an entire cloud, not just TTS.
Open-Source Models (China-Led)
IndexTTS-2 Unique for Video Dubbing
Decouples emotion from timbre and offers millisecond-level duration control — a capability unique to it for lip-sync dubbing — plus Qwen3-tuned natural-language emotion instructions and WER as low as 1.6. For video/short-drama dubbing with second-precise timing, it is the only model that natively delivers.
Pros
- Millisecond duration control (unique)
- Emotion–timbre decoupling
- Qwen3 natural-language emotion instructions
- Apache 2.0, commercial-friendly
Cons
- ~12GB VRAM
- Only Chinese/English + a few languages
- Modest generation speed
Qwen3-TTS Best Open-Source for Commercial
97ms end-to-end latency, 3-second zero-shot cloning, and instruction-level control under a free-for-commercial Apache 2.0 license. With sub-100ms latency and consumer-GPU feasibility, it is one of the lowest-barrier options for self-hosting a real-time voice agent.
Pros
- 97ms end-to-end latency
- 3-second zero-shot cloning
- Apache 2.0, free for commercial use
- Instruction-level control
Cons
- Narrower language coverage than commercial
- Modest high-fidelity ceiling
CosyVoice 3 Top Chinese Open-Source
9 languages + 18 Chinese dialects, ~150ms streaming first chunk, runnable on 8GB VRAM, with Chinese CER among the industry’s lowest. The go-to open-source choice for Chinese self-hosting, forming a “free-for-commercial tier” together with Qwen3-TTS.
Pros
- 9 languages + 18 Chinese dialects
- ~150ms first chunk, streaming-ready
- Runs on 8GB VRAM
- Chinese CER among the lowest
Cons
- High-fidelity ceiling below diffusion-based models
- Fewer languages than top commercial engines
Other Open-Source Models at a Glance
Fish Speech: strong cloning quality, sub-150ms latency, active community — but its open weights cannot be used commercially; you must go through its API.
F5-TTS: MIT-licensed, 330M and lightweight, good flow-matching naturalness, low deployment barrier; cloning similarity is average (SS 0.779).
GPT-SoVITS: MIT-licensed, best few-shot fine-tuning clone fidelity, full VITS toolchain; non-streaming (3-5s), so offline dubbing only.
Kokoro-82M: Apache 2.0, the speed champion (<0.3s), runs on CPU; few languages, no cloning, mid-tier quality.
Higgs Audio V2: Apache 2.0, trained on ~10M hours, strong multi-speaker dialogue expressiveness; resource-hungry and research-oriented.
Moshi: Apache 2.0, the open-source full-duplex benchmark (dual-stream modeling + Inner Monologue); a dialogue model rather than a dubbing engine, mid-tier audio quality.
ChatTTS / VibeVoice: ChatTTS has good conversational naturalness but restricted licensing; VibeVoice supports 90-minute long-form and 4 speakers but was disabled in August 2025 over misuse risk — research use only.
Comparison Matrix: The Whole Field at a Glance
| Platform | Arena ELO | Time-to-First-Audio | Price (per M chars) | Languages | Cloning |
|---|---|---|---|---|---|
| Inworld TTS-1.5 Max | ~1236 | P90 <250ms | Mid-high | 100+ (15 GA) | 5-15s free |
| ElevenLabs Eleven v3 | 1178 | 75ms (Flash) | $120-220 | 32 | 30s instant |
| MiniMax Speech 2.8 HD | 1164 | 400ms+ | $60-100 | 40+ | Zero-shot |
| OpenAI TTS | ~1106 | 200-400ms | $15-30 | 57+ | None |
| Cartesia Sonic 3.5 | ~1054 | 90ms | ~$50 | 40+ | 3s |
| Google Chirp 3 | Top tier | Moderate | $6-30 | 40+ | Limited |
| Volcano Engine Doubao | SuperCLUE #1 (Chinese) | <300ms | ~¥1.3/1K chars (≈$13) | Chinese-first | Yes |
| Alibaba Qwen-Audio-3.0-TTS | SuperCLUE #2 | ~300ms | Bailian tiered | 16 + 20 dialects | 3s |
| iFLYTEK Super-Human | Chinese MOS 4.5+ | Official P50 180ms (measured >1500ms) | Enterprise quote | 20+ dialects | Yes |
Note: ELO figures are a snapshot of the Artificial Analysis blind-test leaderboard (H1 2026); rankings shift weekly, so treat them as ranges, not fixed values. iFLYTEK latency reflects third-party measurements that diverge from official claims — always benchmark before buying.
Scenario Recommendations
Chinese Real-Time Dialogue / Outbound Agent
#1 in Chinese on SuperCLUE + <300ms first chunk + ~¥1.3/1K chars. The overall default for Chinese developers.
English / Multilingual Real-Time Agent
TTFA of 90ms with an SSM architecture that scales cost-effectively. Alternatives: Inworld, ElevenLabs Flash.
Audiobooks / High-Quality Dubbing
#1 in cloning quality and voice library; the quality premium is recoverable through content monetization. Alternative: MiniMax Speech 2.8 HD.
Multilingual Content Going Global
Rich paralinguistics + 40+ languages + word-level timestamps; fits production workflows well.
Video / Short-Drama Dubbing (Lip-Sync, Timing)
The only Apache 2.0 model with millisecond duration control — the only one that natively handles lip-sync and second-precise timing.
Education / Government / Finance (Channel & Compliance)
Most dialects + government/enterprise credentials + MRCP legacy compatibility. Alternatives: Tencent Cloud / Baidu Cloud.
Data-Sensitive / Cost-Sensitive Self-Hosting
Apache 2.0, free for commercial use, runnable on 8GB VRAM. Lightweight alternative: F5-TTS.
Near-Zero-Budget Prototyping
Runs on CPU, speaks in 0.3 seconds. The cheapest tier beyond OpenAI’s free credits.
Decision Framework
1. Chinese real-time dialogue? → Volcano Engine Doubao Voice
2. English / multilingual real-time agent? → Cartesia Sonic 3.5
3. Highest possible quality? → ElevenLabs / Inworld
4. Multilingual content going global? → MiniMax Speech 2.8
5. Video dubbing with timing / lip-sync? → IndexTTS-2
6. Education / government / finance (compliance-heavy)? → iFLYTEK
7. Self-hosting to control cost? → Qwen3-TTS / CosyVoice 3
8. Rock-bottom cost? → OpenAI / Google Flash
Technology Trends
1. “Low-frame-rate tokenizer + LLM semantic tokens + Flow Matching acoustics” is becoming the mainstream paradigm. Baichuan-Audio, Qwen-Audio-3.0, and CosyVoice 2 have all converged on this route — discrete tokens for understanding, continuous representations for high-fidelity synthesis.
2. Full-duplex real-time dialogue is the hottest battleground. Moshi’s dual-stream modeling is the open-source reference architecture; GPT-4o Realtime, Doubao Realtime, and Gemini Live are pushing end-to-end latency into the 300ms range on the closed side. Turn-taking and overlapping-speech modeling are the core challenges.
3. Emotion–timbre decoupling is becoming standard. IndexTTS-2 uses a gradient-reversal layer to separate “who is speaking” from “what emotion,” paired with natural-language emotion instructions; SSML is being replaced by natural-language control.
4. Audio watermarking and embedded compliance. Directions such as in-codec watermarking reflect regulation moving upstream — speech generation is shifting from a “capability problem” to a “traceability problem.”
Compliance & Legal Risk: Two Tracks, China and the U.S.
China: Under the Interim Measures for the Management of Generative AI Services and the Provisions on the Administration of Deep Synthesis of Internet Information Services, TTS-generated content requires visible/invisible labeling, public-facing services must complete filing, and voice cloning requires authorization from the person being cloned. When using domestic commercial APIs, platform-side compliance is clear; when self-hosting open-source models, filing and labeling obligations fall on the user.
United States: In May 2026, leaders such as ElevenLabs faced a BIPA class action in Illinois (alleging training data was used without voice-owner consent); by spring 2026, 46 U.S. states had enacted deepfake-related laws. When procuring overseas TTS services, assess vendors’ training-data compliance statements and indemnification terms.
Key Takeaways
1. Five fronts, no all-round champion. Quality, latency, cost, controllability, and compliance each form their own front line. Choose by real scenario, not by who tops the leaderboard.
2. The quality gap has largely closed. The blind-test top five span the U.S. (Inworld, ElevenLabs, Google) and China (MiniMax) — the leaders are now within the noise range of one another.
3. Latency is a hard metric. The full-chain budget is ~450ms (ASR 100 + LLM first token 150 + TTS first audio 120 + buffer 40); a TTS component over 500ms is unusable in real-time dialogue.
4. Open source has risen — China leads. The Qwen3-TTS / CosyVoice 3 / IndexTTS-2 trio “Apache 2.0 + sub-100ms + consumer GPU” is now the world’s strongest tier.
5. Licenses are a hidden trap. Fish Speech and ChatTTS open weights are not commercially usable, and VibeVoice has been disabled — such traps are more fatal than quality gaps. Verify the license chain before signing.
6. Compliance is a long-term variable. China’s deep-synthesis labeling obligations and the U.S. BIPA litigation plus 46-state deepfake laws will reshape the boundaries of cloning businesses.
© 2026 DeepForgeHub Research. Data sourced from Artificial Analysis, SuperCLUE, IndexTTS technical reports, and public benchmarks and hands-on testing. Model versions and pricing as of September 2026.

