Voice Cloning Models — Global Comparison
Zero-Shot Cloning · Fine-Tuned Cloning · Voice Conversion — Open & Closed — September 2026
Similarity · Naturalness · Cross-Lingual · Sample Efficiency · Real-Time · Cost · Licensing
Executive Summary
Voice cloning entered the “seconds-to-clone, hard-to-distinguish” era in 2026. On the closed side, Cartesia Sonic 3.6 tops the Artificial Analysis Controlled Voice Arena (Elo 1143) with just 3 seconds of reference audio; ElevenLabs v3 pushed expressiveness into the inline-audio-tag era ([whispers], [laughs], 70+ languages); MiniMax Speech 2.8 cut cloning to a $1.50-per-voice one-time fee. On the open side, an open-weight model beat ElevenLabs head-to-head in a blind test for the first time — Resemble AI’s MIT-licensed Chatterbox-Turbo won 65.3% vs 24.5% — while Alibaba’s CosyVoice 2 (Apache-2.0) packaged zero-shot cloning, inline emotion tags, and 150ms streaming first-packet into one complete open stack.
The one-sentence conclusion: closed models still own “quality ceiling + convenience” (peak similarity, expressiveness, ecosystem, compliance packaging); open models own “cost floor + data sovereignty” (free, local, clean licenses). By 2026 the blind-test quality gap has narrowed to where ordinary listeners can’t reliably tell — the deciding factors are no longer “which sounds closer” but “can the voiceprint leave my machine, and what does a million characters cost at scale.”
Key findings:
- “Arena #1” changes hands frequently now. On Artificial Analysis’s Controlled Voice Arena (unified 8-voice, English protocol), Cartesia Sonic 3.6 (1143) leads ElevenLabs v3 (1069) by ~74 Elo — ElevenLabs is no longer the default #1. Open-weight Fish Audio S2.1 Pro / S1 Mini (1011/1010) now sit alongside ElevenLabs Multilingual v2 (1008).
- The cloning threshold is 3–10 seconds. Cartesia 3s, Chatterbox/F5-TTS 10s, MiniMax 5–10s, ElevenLabs IVC 1 minute — fine-tuned cloning’s “30 minutes of material” survives only as the highest-fidelity option (ElevenLabs PVC).
- The price war has hit the floor. API pricing spans 15x from ElevenLabs v3’s $100/1M characters to Speechify Simba 3.0’s $6.6/1M; self-hosted open models cost little more than electricity; MiniMax broke subscription lock-in with a $1.50 one-time clone fee.
- License traps deserve more attention than quality gaps. XTTS-v2 (CPML, no commercial use), IndexTTS-2 (non-commercial without contacting Bilibili), early Fish Speech weights (CC-BY-NC-SA) — “runs fine” is not “sells fine.” Commercially clean: CosyVoice (Apache-2.0), F5-TTS (MIT), Chatterbox (MIT), GPT-SoVITS (MIT), VibeVoice (MIT).
- Compliance tightened across the board. ElevenLabs faces a BIPA class action filed by seven journalists (May 2026); deepfake-enabled vishing grew over 1,600% year-over-year with ~$680K average enterprise loss per attack; EU AI Act Article 50 and China’s deep-synthesis labeling rules add parallel pressure — consent chains and watermarking moved from “nice to have” to “admission ticket.”
1. First, the Fundamentals: Three Technical Roadmaps
“Voice Clone” is one button on product pages but three different things technically. Step one of selection: know which one you need.
| Roadmap | Mechanism | Audio Needed | Similarity Ceiling | Typical Representatives |
|---|---|---|---|---|
| Zero-shot cloning | Reference audio conditions a pretrained model that “imitates” the timbre directly | 3s – 1 min | High (85–95%, hard to catch in blind tests) | Cartesia Sonic, CosyVoice 2, F5-TTS, Chatterbox, MiniMax |
| Fine-tuned / Professional cloning | Per-speaker fine-tuning produces a dedicated acoustic model | 30 min – 3 hrs | Highest (near-indistinguishable) | ElevenLabs PVC, Azure Custom Neural Voice, GPT-SoVITS (self-training) |
| Voice conversion | Replaces the timbre of existing speech, keeping rhythm and timing | Seconds – minutes of reference | High (but depends on source audio quality) | OpenVoice v2, RVC, so-vits-svc, Respeecher (film-grade) |
A common mismatch: wanting to swap language/timbre on existing voiceover (localization) but picking zero-shot TTS — it re-“reads” the text and loses the original performance’s pacing and emotion; the right tools are voice conversion or a dubbing pipeline. Conversely, for generating new narration from text, voice conversion can’t help. Video translation/dubbing (Voice Clone + TTS + lip sync) is zero-shot cloning’s home turf.
2. The Open-Source Camp: Global Representative Solutions
Open-source voice cloning completed the jump from “runs” to “competes” in 2025–2026: zero-shot cloning, emotion control, streaming output, multilingual coverage — the closed camp’s feature forms now mostly exist in open weights, with the remaining gap at the highest fidelity tier and in engineering conveniences.
2.1 Camp Overview
| Solution | Org | Roadmap | Min. Sample | Arena Elo* | License |
|---|---|---|---|---|---|
| CosyVoice 2 / 3 | Alibaba FunAudioLLM (CN) | Zero-shot + emotion tags + streaming | 3–10s | Not listed | Apache-2.0 |
| F5-TTS | SWivid / SJTU (CN) | Zero-shot (Flow Matching) | 10s | Not listed | MIT |
| Chatterbox / Turbo | Resemble AI (US) | Zero-shot + exaggeration control | 5–10s | 937 | MIT |
| IndexTTS-2 / 2.5 | Bilibili (CN) | Zero-shot + timbre/emotion disentanglement | 5–10s | Not listed | Non-commercial without contact |
| XTTS-v2 | Coqui (company wound down; weights live on) | Zero-shot (GPT-style) | 6s | 839 | CPML, no commercial use |
| OpenVoice v2 | MyShell (US/SG) | Voice conversion (timbre swap) | Seconds | 787 | MIT |
| GPT-SoVITS | RVC-Boss community (CN) | Few-shot fine-tuning | 1 min (fine-tune) | Not listed | MIT |
| VibeVoice | Microsoft (US) | Zero-shot long-form (90 min / 4 speakers) | Seconds–tens of seconds | Not listed | MIT |
| Higgs Audio V2/V3 | Boson AI (US) | Zero-shot + conversational | ~10s | 963 (V3) | Open weights |
| MegaTTS3 | ByteDance (CN) | Zero-shot (WavVAE restricted) | Seconds | Not listed | Code Apache / components restricted |
| VoxCPM 1.5 / 2 | OpenBMB / Tsinghua (CN) | Zero-shot, 44.1kHz | ~20s | Not listed | Apache-2.0 |
| RVC | RVC-Project community | Voice conversion (singing/real-time) | ~10 min | Not listed | MIT |
* Elo from the Artificial Analysis Controlled Voice Arena (Aug 2026 board, unified 8-voice English protocol). “Not listed” means no same-protocol public data.
2.2 Key Model Deep-Dives
CosyVoice 2 / 3 Open Source · Apache-2.0Open All-Rounder8.2/10
The highest combined scorer in this report: the closed camp’s full feature form — zero-shot cloning, inline emotion tags ([happy]/[sad]/[angry]), low-latency streaming — open-sourced under Apache-2.0. Top-tier Chinese quality, “usable but not top” English; the 150ms first packet makes it viable for real-time voice agents, which is nearly unique in the open camp.
Pros
- Most complete feature form: cloning + emotion + streaming
- Apache-2.0, zero commercial barriers; large CN community
- 150ms first packet — genuinely usable in live agents
- Self-hosted marginal cost ≈ electricity; data stays local
Cons
- English quality below the closed first tier (docs lean Chinese)
- CosyVoice 3 CPU inference unstable on some platforms
- No arena submission — lacks third-party same-protocol data
- Heavy deployment dependencies; plumbing is DIY
F5-TTS Open Source · MITFastest to Clone7.8/10
The open-source “English quality benchmark” for zero-shot cloning: 10 seconds of reference yields top-tier similarity and naturalness; small model (0.3B), fast inference, mature ComfyUI ecosystem. Language coverage centers on English and Chinese — long-tail languages are the weak spot.
Pros
- High-quality clone from 10s — elite sample efficiency
- MIT + small model: friendly for commercial and edge use
- English cloning quality: open-source first tier
- Mature community (ComfyUI / LoRA fine-tune paths)
Cons
- Narrow language coverage (EN/ZH focus)
- No native emotion-tag expressiveness controls
- Long-form stability is average; segment long texts
- 24kHz sample rate below newer 44.1/48kHz models
Chatterbox / Turbo Open Source · MITBlind-Test Dark Horse7.9/10
One of 2026’s biggest open-source stories: Chatterbox-Turbo beat ElevenLabs in a blind test 65.3% vs 24.5% — the first MIT-licensed open model to overpower the industry default in listening tests (note: vendor-run blind test; read with care). The unique “emotion exaggeration” dial and a 10-minute setup make it the smoothest open dev experience.
Pros
- Blind-test win rate vs ElevenLabs as hard evidence
- Adjustable exaggeration — unique expressiveness control
- MIT license: ship in commercial products directly
- Low deployment barrier (4GB/CPU), streaming output
Cons
- Blind test is vendor-run; arena Elo (937) well below closed leaders
- Long-form stability is average
- Peak fidelity still behind fine-tuned PVC-class cloning
- 23+ languages, but long-tail quality varies
IndexTTS-2 / 2.5 Open Weights · Restricted LicenseEmotion-Control King7.8/10
Among the most acclaimed open cloners for Chinese: timbre/emotion disentanglement is a genuine signature — the same voice can “perform” anger, joy, whispers, or borrow another speaker’s emotion; Chinese naturalness is open-source first tier. The one hard flaw is licensing: weights are open but non-commercial — commercialization requires a separate agreement with Bilibili.
Pros
- Timbre/emotion disentanglement — strongest open expressiveness control
- Top-tier Chinese (and Asian-language) cloning
- Zero-shot 5–10s; solid long-form stability
- Backed by Bilibili, active iteration (now 2.5)
Cons
- Non-commercial license — commercial products must negotiate first
- Mediocre on non-Asian languages
- No arena submission — no same-protocol data
- 8GB+ VRAM; engineering docs thin
XTTS-v2 Open Weights · CPML Non-Commercial7.3/10
Historically irreplaceable — it put “clone any voice from 6 seconds” on every developer’s machine. 17-language coverage still holds up today. But the Nov 2023 weights have been overtaken across the board (Elo 839, bottom of the arena), and CPML’s commercial ban locks it to prototyping and personal use.
Pros
- Easiest setup; one-click installers; Windows friendly
- 17-language zero-shot — coverage still competitive
- Low resource use (4GB VRAM), fast
- Largest fork/tutorial ecosystem
Cons
- CPML bans commercial use — legal risk in products
- Upstream unmaintained; Elo clearly behind new generation
- Similarity decays noticeably on unseen speakers
- 24kHz sample rate caps audio quality
3. The Closed Camp: Global Representative Solutions
The closed camp’s moat has shifted from “how close it sounds” to three places: peak fidelity (fine-tuned cloning), real-time infrastructure (voice agents), and compliance packaging (consent verification, watermarking, SLAs). Price competition has pushed entry tiers down to $5/month.
3.1 Camp Overview
| Solution | Org | Roadmap | Arena Elo* | API ($/1M chars) | Clone Fee |
|---|---|---|---|---|---|
| Cartesia Sonic 3.6 | Cartesia (US) | Zero-shot + sub-100ms real-time | 1143 (#1) | $49 | Included in plan |
| ElevenLabs v3 / Flash v2.5 | ElevenLabs (US) | Zero-shot IVC + fine-tuned PVC | 1069 (v3) | $100 (v3) / $50 (Flash) | IVC included / PVC top tiers |
| Inworld Realtime TTS-2 | Inworld (US) | Zero-shot + realtime agents | 1132 (#2) | $20.8 | Included in plan |
| MiniMax Speech 2.8 | MiniMax (CN) | Zero-shot (5–10s) | 1030 (HD) | $100 (HD) / $60 (Turbo) | $1.50/voice |
| OpenAudio S1 / S2.1 Pro | Fish Audio (CN/US) | Zero-shot (S1 Mini open) | 1011–1010 | $15 | Included in plan |
| Resemble AI (commercial) | Resemble AI (US) | Rapid cloning + detection/watermarks | 937 (Chatterbox) | $25 | Subscription |
| PlayAI (Play 3.0) | PlayHT (US; acquired by Meta, winding down) | Zero-shot / fine-tune | Not listed | Subscription | Included |
| Azure Custom Neural Voice | Microsoft (US) | Professional fine-tuning (gated access) | Not listed | $16–22 | Gated approval |
| Google Chirp 3 HD / Custom Voice | Google (US) | Preset + limited custom | Not listed | $30 | Limited |
| Respeecher | Respeecher (UA) | Film-grade voice conversion | Not listed | Per project | High sample requirement |
| Speechify (Simba 3.0) | Speechify (US) | Zero-shot | 1024 | $6.6 (among lowest) | Included in plan |
* Same Elo protocol as 2.1. Prices from mid-2026 public sources; they move often — verify before contracting.
3.2 Key Model Deep-Dives
ElevenLabs v3 / Flash v2.5 Closed · Commercial APIExpressiveness & Ecosystem Benchmark8.1/10
The industry’s default benchmark: v3’s inline audio tags ([whispers], [laughs], [sighs]) turned “expressiveness” from a parameter into a text primitive; 70+ languages plus the fullest dubbing workbench and Agents ecosystem. PVC fine-tuned cloning remains the “hardest to distinguish from the real person” ceiling. The costs: price ($100/1M) and a split personality on real-time — agents must drop to Flash at lower quality. A BIPA class action (May 2026) questions IVC’s checkbox self-attestation.
Pros
- Expressiveness ceiling: audio tags + multi-speaker dialogue
- PVC is the near-indistinguishable fidelity ceiling — audiobooks’ first pick
- 70+ languages + dubbing/translation/Agents suite
- Most complete compliance stack (Voice Captcha, watermark detector)
Cons
- Priciest tier: $100/1M; credit complaints (failed generations bill too)
- v3 isn’t real-time — agents must drop to Flash (quality falls)
- Arena #1 lost to Cartesia/Inworld
- BIPA suit; self-attested consent criticized
Cartesia Sonic 3.6 Closed · Commercial APIArena #1 · Real-Time King8.2/10
2026’s biggest arena dark horse: an SSM (state-space model) architecture replaces the mainstream DiT and captures two previously conflicting titles — top cloning quality and 40ms real time. 3-second cloning, $49/1M characters, $5/mo entry. The de-facto first pick for voice-agent stacks.
Pros
- Third-party arena #1 (unified protocol, all players)
- 3s cloning + 40ms latency — unmatched for live interaction
- Efficient SSM architecture; mid pricing ($49/1M)
- Low entry barrier ($5/mo)
Cons
- Language coverage (42 Pro) behind ElevenLabs’ 70+
- Expression-tag system less rich than v3
- No professional fine-tuned (PVC-class) tier
- Offline content ecosystem (audiobooks, dubbing) thinner than ElevenLabs
MiniMax Speech 2.8 Closed · Commercial APIValue King7.8/10
The player that rewrote cloning cost structure: a $1.50 one-time Rapid Voice Cloning fee breaks subscription lock-in; Speech-02 topped two speech arenas at launch; a 200,000-character async request limit suits long-form batch. First-tier Chinese quality — a high-value choice for bilingual (EN/ZH) video dubbing pipelines.
Pros
- $1.50 one-time clone fee — near-zero trial cost
- 200K-char requests; batch long-form friendly
- Balanced Chinese quality and 32-language coverage
- Free tier (10K credits) to validate first
Cons
- HD at $100/1M isn’t cheap — value comes from Turbo and clone fees
- Streaming limited to ≤5,000-character requests
- China-region deployments; extra compliance review for global enterprises
- Expressiveness controls (emotion tags) behind v3/CosyVoice
Inworld Realtime TTS-2 Closed · Commercial API7.9/10
The quiet champion of gaming and AI-companion scenarios: Realtime TTS-2 ranks #2 on the arena at 40% of Cartesia’s price ($20.8/1M), and the Flash tier runs $10.4/1M — “top-two quality + below-median price” is the best value in the agent track. General dubbing/audiobook ecosystems lag ElevenLabs.
Pros
- Arena #2 — third-party quality backing
- Cheapest among the top two (Flash $10.4/1M)
- Deep game-engine/agent-framework integrations
- Strong emotion and roleplay tuning
Cons
- Agent/game focus; offline dubbing pipelines weak
- Moderate language coverage (28)
- Basic clone-management features
- Lower brand awareness; fewer community resources
OpenAudio S1 / S2.1 Pro (Fish Audio) Closed API + Open Mini7.6/10
A hybrid “half open, half closed” play: S1 Mini ships open weights (self-hostable), while commercial S2.1 Pro ties ElevenLabs’ older flagship on the arena. Solid EN/ZH/JA coverage with native emotion markers (e.g., (angry)) and multi-speaker dialogue. A natural “try open first, switch to API at scale” path.
Pros
- Dual track: open Mini + commercial API, frictionless migration
- Open-weight first tier on the arena (nears ElevenLabs’ old flagship)
- Friendly $15/1M; emotion markers and multi-speaker built in
- Clones from 10–30s
Cons
- Open Mini weights not fully commercial-free (CC-BY-NC family)
- Enterprise compliance tooling (consent/watermarking) behind ElevenLabs
- Few long-tail languages
- Small company — SLA and stability to watch
4. Head-to-Head: Open vs Closed, Dimension by Dimension
| Dimension | Open-Source Camp | Closed Camp | Edge |
|---|---|---|---|
| Similarity (top tier) | Zero-shot reaches the “hard to catch in blind tests” bar, but lacks a fine-tuning tier | PVC-class fine-tuned cloning remains the near-indistinguishable ceiling | Closed |
| Naturalness / expressiveness | CosyVoice emotion tags and IndexTTS disentanglement match the feature forms | v3 audio tags + arena-leading quality | Closed, slightly |
| Cloning efficiency | F5-TTS 10s, Chatterbox 5–10s | Cartesia 3s, MiniMax 5–10s | Even (Cartesia’s extreme lead) |
| Cross-lingual | XTTS 17, Chatterbox 23; long-tail quality varies | ElevenLabs 70+, Google 380+ preset voices | Closed |
| Real-time / streaming | CosyVoice 150ms first packet (open’s only), Chatterbox streaming | Cartesia 40ms, ElevenLabs Flash 75ms | Closed (extremes); open is usable |
| Cost | Self-host ≈ electricity; no metering | $6.6–100/1M; MiniMax $1.50/voice lowest entry | Open (at scale) |
| License & data sovereignty | MIT/Apache: commercial, private, local — voiceprints stay in-house | Data to cloud; compliance packaging (watermarks/verification) is the counterweight | Open (weights), closed (tooling) |
| Engineering convenience | Deployment, batching, monitoring all DIY | API-ready, SLAs, consoles, integration ecosystems | Closed |
Closed wins five of eight, open wins three — but as in the digital-human report, the three open wins (cost, license, sovereignty) are precisely the most expensive items for scaled commercial deployment.
5. Scoring & Head-to-Head Evaluation
Scoring disclaimer: the scores and star ratings below are this report’s unofficial evaluation, compiled from public materials, the Artificial Analysis Controlled Voice Arena, vendor documentation, and community testing. They are not official benchmark numbers. Arena Elo uses a unified English protocol; multilingual capability is assessed from public specs and community feedback.
5.1 Criteria & Weights (customized for the voice-cloning task)
| Criterion | Weight | What it measures |
|---|---|---|
| Voice similarity | 25% | How distinguishable the clone is from the real speaker (zero-shot and fine-tuned assessed accordingly) |
| Naturalness & expressiveness | 20% | Listening quality, emotion/style controllability |
| Cross-lingual | 15% | Language coverage and cross-language clone quality — a hard requirement for video localization |
| Sample efficiency | 10% | Reference-audio duration and onboarding cost |
| Speed & real-time | 10% | RTF, first-packet latency, streaming |
| Cost | 10% | API pricing / self-host marginal cost / clone-fee structure |
| License & compliance | 10% | Commercial license permissiveness + consent/watermark tooling |
5.2 Overall Scoreboard
| Model | Camp | Arena Elo* | Score | Stars | Progress |
|---|---|---|---|---|---|
| Cartesia Sonic 3.6 | Closed | 1143 | 8.2 | ★★★★★ | |
| CosyVoice 2 / 3 | Open | Not listed | 8.2 | ★★★★★ | |
| ElevenLabs v3 | Closed | 1069 | 8.1 | ★★★★☆ | |
| Chatterbox / Turbo | Open | 937 | 7.9 | ★★★★☆ | |
| Inworld Realtime TTS-2 | Closed | 1132 | 7.9 | ★★★★☆ | |
| MiniMax Speech 2.8 | Closed | 1030 | 7.8 | ★★★★☆ | |
| F5-TTS | Open | Not listed | 7.8 | ★★★★☆ | |
| IndexTTS-2 | Open (restricted license) | Not listed | 7.8 | ★★★★☆ | |
| OpenAudio S1 / S2.1 Pro | Hybrid | 1011 | 7.6 | ★★★★☆ | |
| VibeVoice | Open | Not listed | 7.5 | ★★★★★ | |
| Higgs Audio V3 | Open | 963 | 7.5 | ★★★★★ | |
| OpenVoice v2 | Open | 787 | 7.4 | ★★★★★ | |
| XTTS-v2 | Open (non-commercial) | 839 | 7.3 | ★★★★★ | |
| GPT-SoVITS | Open | Not listed | 7.3 | ★★★★★ |
* “Not listed” means no same-protocol arena data — it is not a quality ranking. The composite includes cost and licensing, so it deliberately differs from a pure-quality Elo order — e.g., ElevenLabs ranks higher on pure quality than CosyVoice, but CosyVoice’s free price and Apache license flip the composite. Differences within ±0.3 are ties.
5.3 Seven-Criterion Detail
Legend: ✓ strength · ◐ average · ✗ weakness
| Model | Similarity 25% |
Naturalness 20% |
Cross-Lingual 15% |
Efficiency 10% |
Speed/RT 10% |
Cost 10% |
License 10% |
|---|---|---|---|---|---|---|---|
| Cartesia Sonic 3.6 | 9.0 | 9.0 | 7.5 | 9.5 | 10 | 6.0 | 5.0 |
| CosyVoice 2 | 7.5 | 7.5 | 8.0 | 8.5 | 8.5 | 10 | 9.5 |
| ElevenLabs v3 | 9.5 | 9.5 | 9.5 | 8.0 | 6.0 | 4.0 | 6.0 |
| Chatterbox | 7.5 | 8.0 | 7.0 | 8.5 | 8.0 | 8.5 | 9.0 |
| Inworld TTS-2 | 8.5 | 8.5 | 7.0 | 8.0 | 9.0 | 8.5 | 5.0 |
| MiniMax Speech 2.8 | 8.5 | 8.5 | 8.0 | 8.5 | 7.5 | 8.0 | 4.0 |
| F5-TTS | 7.5 | 7.5 | 6.0 | 9.0 | 7.0 | 10 | 9.0 |
| IndexTTS-2 | 8.5 | 8.5 | 6.5 | 8.5 | 7.0 | 10 | 4.5 |
| OpenAudio S1/S2 | 8.0 | 8.0 | 7.5 | 8.0 | 6.5 | 9.0 | 5.5 |
| VibeVoice | 7.5 | 7.5 | 6.0 | 8.0 | 5.0 | 10 | 9.5 |
| Higgs Audio V3 | 7.0 | 7.5 | 7.0 | 7.5 | 6.5 | 10 | 8.0 |
| OpenVoice v2 | 6.5 | 6.0 | 6.0 | 9.0 | 9.0 | 10 | 9.0 |
| XTTS-v2 | 6.5 | 6.5 | 8.0 | 8.5 | 8.0 | 10 | 5.0 |
| GPT-SoVITS | 8.0 | 7.0 | 5.0 | 6.0 | 6.5 | 10 | 9.0 |
5.4 Category Champions
Audio-tag expressiveness + the near-indistinguishable fine-tuned ceiling.
#1 on a unified protocol across all players — and also the lowest latency.
The most feature-complete open stack: cloning + emotion + streaming.
Closed extreme at Cartesia; the practical open line at F5-TTS.
A’s voice with B’s emotion — a unique open trick (mind the license).
One-time clone fee + Turbo $60/1M — the cheapest closed path at scale.
Fine-tune on 1 minute for top Chinese similarity (hands-on training).
Podcast/audiobook-length multi-speaker audio, MIT-licensed.
5.5 Head-to-Head Verdicts
① Multilingual Video Dubbing: ElevenLabs v3 vs CosyVoice 2
Video localization demands 30+ languages, expressiveness, and batch stability. v3’s 70+ languages, audio tags, and dubbing workbench are turnkey; CosyVoice is free but English tuning and plumbing are on you. Under ~500K chars/month, ElevenLabs saves ops; at scale, self-hosted CosyVoice wins by an order of magnitude. Verdict: ElevenLabs at low volume, CosyVoice at high volume.
② Real-Time Voice Agents: Cartesia Sonic 3.6 vs MiniMax 2.8
For agents, latency is the experience. Cartesia’s 40ms Turbo plus arena #1 and mature agent-stack integrations beat MiniMax, whose streaming caps at 5,000 characters. Winner: Cartesia Sonic 3.6 — budget alternative: Inworld Flash ($10.4/1M).
③ Local Zero-Shot Cloning: F5-TTS vs XTTS-v2
Both classics on a single RTX 4090. F5-TTS wins quality, sample efficiency (10s), and an MIT license; XTTS-v2 wins setup convenience and 17 languages, but CPML’s commercial ban is a hard flaw. Winner: F5-TTS — keep XTTS-v2 for prototypes and personal use only.
④ High-Similarity Chinese Cloning: IndexTTS-2 vs GPT-SoVITS
IndexTTS-2 is zero-shot with emotion disentanglement, usable in minutes; GPT-SoVITS reaches higher similarity after fine-tuning with 1 minute of data and ships MIT-clean. Verdict: speed and emotion → IndexTTS-2 (negotiate commercial terms first); clean licensing and peak similarity → GPT-SoVITS.
6. Scenario Recommendations
Video Translation / Multilingual Dubbing
Pair with lip-sync pipelines for 30+ language localization; start on ElevenLabs’ ecosystem, graduate to MiniMax or self-hosted CosyVoice at scale.
Real-Time Voice Agents / Support Bots
40ms latency plus arena-#1 quality; budget pick: Inworld Flash ($10.4/1M).
Audiobooks / Podcast Batch Production
For “hardest to distinguish,” use ElevenLabs professional cloning; for budget multi-speaker long-form, VibeVoice (90 min / 4 speakers, MIT) locally.
VTubing / Game Characters
Emotion-disentangled performance plus fine-tuned similarity; resolve IndexTTS-2 licensing before commercial use, or stay all-MIT with GPT-SoVITS/RVC.
Timbre Swap on Existing Audio / AI Covers
The voice-conversion route: keep the performance, swap the timbre; RVC for singing, OpenVoice v2 for speech — both MIT.
Privacy-Sensitive / On-Prem
A fully local pipeline (Apache/MIT) — voiceprints never leave the machine. The only acceptable answer for finance, healthcare, and government.
7. Decision Framework
The Five-Question Method
Q1: Re-speak or re-voice? Generating new narration → zero-shot/fine-tuned cloning (TTS route). Keeping a performance but swapping timbre → voice conversion (OpenVoice/RVC/Respeecher).
Q2: Can data leave the machine? Voiceprints are biometric data. If not → open self-hosting (CosyVoice/F5-TTS/Chatterbox). If yes → continue.
Q3: Real time? Live agents/streaming → Cartesia (40ms) or Inworld; offline content → v3/MiniMax/CosyVoice batch.
Q4: What volume? < 500K chars/month → subscriptions ($5–22/mo); millions+ → API negotiation or open self-hosting; MiniMax’s $1.50/voice is the lowest trial threshold.
Q5: Is the license clean? Before launch check three things: model license (XTTS’s CPML and IndexTTS-2’s non-commercial terms are the classic traps), the consent chain (written permission from the voice owner), and labeling duties (AI-generated content disclosure).
8. Trends
- Architectural diversification: SSMs challenge DiTs. Cartesia’s state-space models took both quality #1 and 40ms real time, breaking the old “diffusion = slow, autoregressive = fast” rule; flow matching (F5-TTS) is the new open-source default.
- Expressiveness is the new battlefield. ElevenLabs v3’s audio tags, CosyVoice’s emotion tags, IndexTTS-2’s timbre/emotion disentanglement, Chatterbox’s exaggeration dial — the competition moved from “sounding like” to “performing like.”
- Open weights beat closed in blind tests for the first time. Chatterbox-Turbo’s 65.3% over ElevenLabs (vendor-run) plus open S1 Mini tying ElevenLabs’ old flagship on the arena — the open “good enough” line has crossed most commercial scenarios.
- Prices keep breaking floors. ElevenLabs cut up to 55% in May 2026; Speechify at $6.6/1M, Inworld Flash $10.4/1M, Fish $15/1M — the per-character floor fell an order of magnitude within a year, blurring subscription vs metered boundaries.
- From single models to pipelines. Leading products embed cloning into larger pipelines (dubbing + lip sync + translation; agents + RAG) — single-API differentiation is compressing; latency, concurrency, and compliance become the new selling points.
9. Compliance & Ethics Notes (Voice Cloning’s Required Reading)
- Voiceprints are biometric data. Cloning someone’s voice requires explicit consent — ElevenLabs faced a BIPA class action (May 2026) over IVC’s checkbox self-attestation, which Consumer Reports called “no meaningful technical barrier.” Commercial products should use recorded-statement verification (Voice-Captcha-style).
- Fraud is the biggest negative externality. Deepfake-driven vishing grew over 1,600% year-over-year, enterprises lose ~$680K per attack on average, and the Arup case cost $25.6M in one call. The stronger the cloning capability, the more it needs paired watermarking and detection (ElevenLabs AI Speech Classifier, Resemble’s detector).
- Labeling duties cut both ways. EU AI Act Article 50 (effective Aug 2026) requires prominent AI-content disclosure; China’s deep-synthesis rules mandate explicit and implicit labels. Embed hidden watermarks and add AI labels to cloned-voice output.
- Commercial license red lines. XTTS-v2 (CPML), IndexTTS-2 (non-commercial), parts of Fish Speech weights (CC-BY-NC) — open ≠ commercially free. Verify licenses and voice-data provenance model by model before launch.
10. Key Takeaways
- “Arena #1” changes hands — ElevenLabs’ default advantage is gone. Cartesia Sonic 3.6 (1143) and Inworld TTS-2 (1132) lead ElevenLabs v3 (1069) on a unified third-party protocol. Select by testing your own material, not by brand default.
- The open-vs-closed quality gap is now inside blind-test noise. Chatterbox’s blind-test win (vendor-run) and open S1 Mini tying ElevenLabs’ old flagship mean the closed camp’s hard advantages are down to the PVC fidelity ceiling, 70+ languages, and compliance tooling.
- A new cost structure has appeared. MiniMax’s $1.50 one-time clone fee breaks subscription lock-in; self-hosted marginal cost approaches zero — at scale, voice is the cheapest media asset to duplicate.
- This is the license-trap era. Non-commercial licenses (XTTS, IndexTTS-2) and compliance litigation (ElevenLabs BIPA) arrived together — the first 2026 selection question isn’t quality, it’s “can I legally sell this voice.”
- The three roadmaps don’t substitute for each other. Localization needs voice conversion to preserve performance; new content needs zero-shot for iteration speed; IP-level voice assets need fine-tuning for the ceiling. Pick the route before comparing models.
- Chinese teams anchor both the open and value ends. CosyVoice (Apache), F5-TTS (MIT), IndexTTS-2, GPT-SoVITS, Fish Audio, MiniMax — the world’s open voice-cloning weights and lowest price bands are supplied largely by Chinese teams.
Run Voice Cloning on Your Own Machine
DeepVideo is a locally running AI video translation and dubbing client: Voice Clone in 30+ languages, 100+ preset voices, TTS + lip sync — voiceprint data is processed on your device, never uploaded to the cloud.
Try DeepVideo Free
Realistic AI Video Dubbing at Just $0.17/min
High quality · Low price · Local client · Security
Free tier: 18 minutes total + 2 minutes daily · Local processing, no cloud upload
DeepForgeHub Research · Voice Cloning Models Global Comparison · September 2026
deepforgehub.com · Compiled from public sources; scores are unofficial evaluations

