Voice Cloning Models Compared 2026: Open Source vs Closed
Voice cloning model comparison table: similarity, naturalness, cross-lingual, efficiency and speed scores for leading open and closed TTS models

Voice Cloning Models: Global Comparison Report (2026)

Voice Cloning Models — Global Comparison

Zero-Shot Cloning · Fine-Tuned Cloning · Voice Conversion — Open & Closed — September 2026

Similarity · Naturalness · Cross-Lingual · Sample Efficiency · Real-Time · Cost · Licensing

Executive Summary

Voice cloning entered the “seconds-to-clone, hard-to-distinguish” era in 2026. On the closed side, Cartesia Sonic 3.6 tops the Artificial Analysis Controlled Voice Arena (Elo 1143) with just 3 seconds of reference audio; ElevenLabs v3 pushed expressiveness into the inline-audio-tag era ([whispers], [laughs], 70+ languages); MiniMax Speech 2.8 cut cloning to a $1.50-per-voice one-time fee. On the open side, an open-weight model beat ElevenLabs head-to-head in a blind test for the first time — Resemble AI’s MIT-licensed Chatterbox-Turbo won 65.3% vs 24.5% — while Alibaba’s CosyVoice 2 (Apache-2.0) packaged zero-shot cloning, inline emotion tags, and 150ms streaming first-packet into one complete open stack.

The one-sentence conclusion: closed models still own “quality ceiling + convenience” (peak similarity, expressiveness, ecosystem, compliance packaging); open models own “cost floor + data sovereignty” (free, local, clean licenses). By 2026 the blind-test quality gap has narrowed to where ordinary listeners can’t reliably tell — the deciding factors are no longer “which sounds closer” but “can the voiceprint leave my machine, and what does a million characters cost at scale.”

Key findings:

  • “Arena #1” changes hands frequently now. On Artificial Analysis’s Controlled Voice Arena (unified 8-voice, English protocol), Cartesia Sonic 3.6 (1143) leads ElevenLabs v3 (1069) by ~74 Elo — ElevenLabs is no longer the default #1. Open-weight Fish Audio S2.1 Pro / S1 Mini (1011/1010) now sit alongside ElevenLabs Multilingual v2 (1008).
  • The cloning threshold is 3–10 seconds. Cartesia 3s, Chatterbox/F5-TTS 10s, MiniMax 5–10s, ElevenLabs IVC 1 minute — fine-tuned cloning’s “30 minutes of material” survives only as the highest-fidelity option (ElevenLabs PVC).
  • The price war has hit the floor. API pricing spans 15x from ElevenLabs v3’s $100/1M characters to Speechify Simba 3.0’s $6.6/1M; self-hosted open models cost little more than electricity; MiniMax broke subscription lock-in with a $1.50 one-time clone fee.
  • License traps deserve more attention than quality gaps. XTTS-v2 (CPML, no commercial use), IndexTTS-2 (non-commercial without contacting Bilibili), early Fish Speech weights (CC-BY-NC-SA) — “runs fine” is not “sells fine.” Commercially clean: CosyVoice (Apache-2.0), F5-TTS (MIT), Chatterbox (MIT), GPT-SoVITS (MIT), VibeVoice (MIT).
  • Compliance tightened across the board. ElevenLabs faces a BIPA class action filed by seven journalists (May 2026); deepfake-enabled vishing grew over 1,600% year-over-year with ~$680K average enterprise loss per attack; EU AI Act Article 50 and China’s deep-synthesis labeling rules add parallel pressure — consent chains and watermarking moved from “nice to have” to “admission ticket.”

1. First, the Fundamentals: Three Technical Roadmaps

“Voice Clone” is one button on product pages but three different things technically. Step one of selection: know which one you need.

Roadmap Mechanism Audio Needed Similarity Ceiling Typical Representatives
Zero-shot cloning Reference audio conditions a pretrained model that “imitates” the timbre directly 3s – 1 min High (85–95%, hard to catch in blind tests) Cartesia Sonic, CosyVoice 2, F5-TTS, Chatterbox, MiniMax
Fine-tuned / Professional cloning Per-speaker fine-tuning produces a dedicated acoustic model 30 min – 3 hrs Highest (near-indistinguishable) ElevenLabs PVC, Azure Custom Neural Voice, GPT-SoVITS (self-training)
Voice conversion Replaces the timbre of existing speech, keeping rhythm and timing Seconds – minutes of reference High (but depends on source audio quality) OpenVoice v2, RVC, so-vits-svc, Respeecher (film-grade)

A common mismatch: wanting to swap language/timbre on existing voiceover (localization) but picking zero-shot TTS — it re-“reads” the text and loses the original performance’s pacing and emotion; the right tools are voice conversion or a dubbing pipeline. Conversely, for generating new narration from text, voice conversion can’t help. Video translation/dubbing (Voice Clone + TTS + lip sync) is zero-shot cloning’s home turf.

2. The Open-Source Camp: Global Representative Solutions

Open-source voice cloning completed the jump from “runs” to “competes” in 2025–2026: zero-shot cloning, emotion control, streaming output, multilingual coverage — the closed camp’s feature forms now mostly exist in open weights, with the remaining gap at the highest fidelity tier and in engineering conveniences.

2.1 Camp Overview

Solution Org Roadmap Min. Sample Arena Elo* License
CosyVoice 2 / 3 Alibaba FunAudioLLM (CN) Zero-shot + emotion tags + streaming 3–10s Not listed Apache-2.0
F5-TTS SWivid / SJTU (CN) Zero-shot (Flow Matching) 10s Not listed MIT
Chatterbox / Turbo Resemble AI (US) Zero-shot + exaggeration control 5–10s 937 MIT
IndexTTS-2 / 2.5 Bilibili (CN) Zero-shot + timbre/emotion disentanglement 5–10s Not listed Non-commercial without contact
XTTS-v2 Coqui (company wound down; weights live on) Zero-shot (GPT-style) 6s 839 CPML, no commercial use
OpenVoice v2 MyShell (US/SG) Voice conversion (timbre swap) Seconds 787 MIT
GPT-SoVITS RVC-Boss community (CN) Few-shot fine-tuning 1 min (fine-tune) Not listed MIT
VibeVoice Microsoft (US) Zero-shot long-form (90 min / 4 speakers) Seconds–tens of seconds Not listed MIT
Higgs Audio V2/V3 Boson AI (US) Zero-shot + conversational ~10s 963 (V3) Open weights
MegaTTS3 ByteDance (CN) Zero-shot (WavVAE restricted) Seconds Not listed Code Apache / components restricted
VoxCPM 1.5 / 2 OpenBMB / Tsinghua (CN) Zero-shot, 44.1kHz ~20s Not listed Apache-2.0
RVC RVC-Project community Voice conversion (singing/real-time) ~10 min Not listed MIT

* Elo from the Artificial Analysis Controlled Voice Arena (Aug 2026 board, unified 8-voice English protocol). “Not listed” means no same-protocol public data.

2.2 Key Model Deep-Dives

CosyVoice 2 / 3 Open Source · Apache-2.0Open All-Rounder8.2/10

Org: Alibaba FunAudioLLM
Focus: Zero-shot cloning + emotion tags + 150ms streaming first packet
Sample: 3–10s reference
VRAM: ~8GB (0.5B)
License: Apache-2.0

The highest combined scorer in this report: the closed camp’s full feature form — zero-shot cloning, inline emotion tags ([happy]/[sad]/[angry]), low-latency streaming — open-sourced under Apache-2.0. Top-tier Chinese quality, “usable but not top” English; the 150ms first packet makes it viable for real-time voice agents, which is nearly unique in the open camp.

Pros

  • Most complete feature form: cloning + emotion + streaming
  • Apache-2.0, zero commercial barriers; large CN community
  • 150ms first packet — genuinely usable in live agents
  • Self-hosted marginal cost ≈ electricity; data stays local

Cons

  • English quality below the closed first tier (docs lean Chinese)
  • CosyVoice 3 CPU inference unstable on some platforms
  • No arena submission — lacks third-party same-protocol data
  • Heavy deployment dependencies; plumbing is DIY

F5-TTS Open Source · MITFastest to Clone7.8/10

Org: SWivid / Shanghai Jiao Tong University
Focus: Flow-matching zero-shot cloning with 10s reference
Params: 0.3B, 24kHz
Speed: ~4s per sentence (RTX 3060 class)

The open-source “English quality benchmark” for zero-shot cloning: 10 seconds of reference yields top-tier similarity and naturalness; small model (0.3B), fast inference, mature ComfyUI ecosystem. Language coverage centers on English and Chinese — long-tail languages are the weak spot.

Pros

  • High-quality clone from 10s — elite sample efficiency
  • MIT + small model: friendly for commercial and edge use
  • English cloning quality: open-source first tier
  • Mature community (ComfyUI / LoRA fine-tune paths)

Cons

  • Narrow language coverage (EN/ZH focus)
  • No native emotion-tag expressiveness controls
  • Long-form stability is average; segment long texts
  • 24kHz sample rate below newer 44.1/48kHz models

Chatterbox / Turbo Open Source · MITBlind-Test Dark Horse7.9/10

Org: Resemble AI (commercial company, open model)
Focus: Zero-shot cloning + 0–100% “exaggeration” dial
Sample: 5–10s
Languages: 23+

One of 2026’s biggest open-source stories: Chatterbox-Turbo beat ElevenLabs in a blind test 65.3% vs 24.5% — the first MIT-licensed open model to overpower the industry default in listening tests (note: vendor-run blind test; read with care). The unique “emotion exaggeration” dial and a 10-minute setup make it the smoothest open dev experience.

Pros

  • Blind-test win rate vs ElevenLabs as hard evidence
  • Adjustable exaggeration — unique expressiveness control
  • MIT license: ship in commercial products directly
  • Low deployment barrier (4GB/CPU), streaming output

Cons

  • Blind test is vendor-run; arena Elo (937) well below closed leaders
  • Long-form stability is average
  • Peak fidelity still behind fine-tuned PVC-class cloning
  • 23+ languages, but long-tail quality varies

IndexTTS-2 / 2.5 Open Weights · Restricted LicenseEmotion-Control King7.8/10

Org: Bilibili
Focus: Zero-shot cloning + timbre/emotion disentanglement (mix A’s voice with B’s emotion)
Sample: 5–10s
License: Non-commercial without contacting Bilibili

Among the most acclaimed open cloners for Chinese: timbre/emotion disentanglement is a genuine signature — the same voice can “perform” anger, joy, whispers, or borrow another speaker’s emotion; Chinese naturalness is open-source first tier. The one hard flaw is licensing: weights are open but non-commercial — commercialization requires a separate agreement with Bilibili.

Pros

  • Timbre/emotion disentanglement — strongest open expressiveness control
  • Top-tier Chinese (and Asian-language) cloning
  • Zero-shot 5–10s; solid long-form stability
  • Backed by Bilibili, active iteration (now 2.5)

Cons

  • Non-commercial license — commercial products must negotiate first
  • Mediocre on non-Asian languages
  • No arena submission — no same-protocol data
  • 8GB+ VRAM; engineering docs thin

XTTS-v2 Open Weights · CPML Non-Commercial7.3/10

Org: Coqui (company wound down in 2024; weights and forks remain in wide use)
Focus: The original zero-shot cloning popularizer — 17 languages / 6s reference
Speed: ~2s per sentence; runs on 4GB VRAM
License: Coqui Public Model License — no commercial use

Historically irreplaceable — it put “clone any voice from 6 seconds” on every developer’s machine. 17-language coverage still holds up today. But the Nov 2023 weights have been overtaken across the board (Elo 839, bottom of the arena), and CPML’s commercial ban locks it to prototyping and personal use.

Pros

  • Easiest setup; one-click installers; Windows friendly
  • 17-language zero-shot — coverage still competitive
  • Low resource use (4GB VRAM), fast
  • Largest fork/tutorial ecosystem

Cons

  • CPML bans commercial use — legal risk in products
  • Upstream unmaintained; Elo clearly behind new generation
  • Similarity decays noticeably on unseen speakers
  • 24kHz sample rate caps audio quality

3. The Closed Camp: Global Representative Solutions

The closed camp’s moat has shifted from “how close it sounds” to three places: peak fidelity (fine-tuned cloning), real-time infrastructure (voice agents), and compliance packaging (consent verification, watermarking, SLAs). Price competition has pushed entry tiers down to $5/month.

3.1 Camp Overview

Solution Org Roadmap Arena Elo* API ($/1M chars) Clone Fee
Cartesia Sonic 3.6 Cartesia (US) Zero-shot + sub-100ms real-time 1143 (#1) $49 Included in plan
ElevenLabs v3 / Flash v2.5 ElevenLabs (US) Zero-shot IVC + fine-tuned PVC 1069 (v3) $100 (v3) / $50 (Flash) IVC included / PVC top tiers
Inworld Realtime TTS-2 Inworld (US) Zero-shot + realtime agents 1132 (#2) $20.8 Included in plan
MiniMax Speech 2.8 MiniMax (CN) Zero-shot (5–10s) 1030 (HD) $100 (HD) / $60 (Turbo) $1.50/voice
OpenAudio S1 / S2.1 Pro Fish Audio (CN/US) Zero-shot (S1 Mini open) 1011–1010 $15 Included in plan
Resemble AI (commercial) Resemble AI (US) Rapid cloning + detection/watermarks 937 (Chatterbox) $25 Subscription
PlayAI (Play 3.0) PlayHT (US; acquired by Meta, winding down) Zero-shot / fine-tune Not listed Subscription Included
Azure Custom Neural Voice Microsoft (US) Professional fine-tuning (gated access) Not listed $16–22 Gated approval
Google Chirp 3 HD / Custom Voice Google (US) Preset + limited custom Not listed $30 Limited
Respeecher Respeecher (UA) Film-grade voice conversion Not listed Per project High sample requirement
Speechify (Simba 3.0) Speechify (US) Zero-shot 1024 $6.6 (among lowest) Included in plan

* Same Elo protocol as 2.1. Prices from mid-2026 public sources; they move often — verify before contracting.

3.2 Key Model Deep-Dives

ElevenLabs v3 / Flash v2.5 Closed · Commercial APIExpressiveness & Ecosystem Benchmark8.1/10

Org: ElevenLabs ($11B valuation Series D, Feb 2026; $500M ARR by May)
Focus: Instant cloning IVC (1 min) / Professional PVC (30 min–3 hrs)
Languages: 70+ (v3), 32 (Flash)
Speed: v3 not real-time (5,000 chars/request); Flash TTFB ~75ms

The industry’s default benchmark: v3’s inline audio tags ([whispers], [laughs], [sighs]) turned “expressiveness” from a parameter into a text primitive; 70+ languages plus the fullest dubbing workbench and Agents ecosystem. PVC fine-tuned cloning remains the “hardest to distinguish from the real person” ceiling. The costs: price ($100/1M) and a split personality on real-time — agents must drop to Flash at lower quality. A BIPA class action (May 2026) questions IVC’s checkbox self-attestation.

Pros

  • Expressiveness ceiling: audio tags + multi-speaker dialogue
  • PVC is the near-indistinguishable fidelity ceiling — audiobooks’ first pick
  • 70+ languages + dubbing/translation/Agents suite
  • Most complete compliance stack (Voice Captcha, watermark detector)

Cons

  • Priciest tier: $100/1M; credit complaints (failed generations bill too)
  • v3 isn’t real-time — agents must drop to Flash (quality falls)
  • Arena #1 lost to Cartesia/Inworld
  • BIPA suit; self-attested consent criticized

Cartesia Sonic 3.6 Closed · Commercial APIArena #1 · Real-Time King8.2/10

Org: Cartesia (US)
Focus: 3-second zero-shot cloning + sub-100ms voice agents
Arena: Artificial Analysis Controlled Voice #1 (Elo 1143)
Speed: Sonic 3.5 ~90ms; Turbo tier 40ms
Source: cartesia.ai

2026’s biggest arena dark horse: an SSM (state-space model) architecture replaces the mainstream DiT and captures two previously conflicting titles — top cloning quality and 40ms real time. 3-second cloning, $49/1M characters, $5/mo entry. The de-facto first pick for voice-agent stacks.

Pros

  • Third-party arena #1 (unified protocol, all players)
  • 3s cloning + 40ms latency — unmatched for live interaction
  • Efficient SSM architecture; mid pricing ($49/1M)
  • Low entry barrier ($5/mo)

Cons

  • Language coverage (42 Pro) behind ElevenLabs’ 70+
  • Expression-tag system less rich than v3
  • No professional fine-tuned (PVC-class) tier
  • Offline content ecosystem (audiobooks, dubbing) thinner than ElevenLabs

MiniMax Speech 2.8 Closed · Commercial APIValue King7.8/10

Org: MiniMax (CN)
Focus: 5–10s rapid cloning; HD/Turbo tiers
Pricing: HD $100 / Turbo $60 per 1M chars; cloning $1.50/voice one-time
Languages: 32 (Chinese-strong)
Source: minimax.io

The player that rewrote cloning cost structure: a $1.50 one-time Rapid Voice Cloning fee breaks subscription lock-in; Speech-02 topped two speech arenas at launch; a 200,000-character async request limit suits long-form batch. First-tier Chinese quality — a high-value choice for bilingual (EN/ZH) video dubbing pipelines.

Pros

  • $1.50 one-time clone fee — near-zero trial cost
  • 200K-char requests; batch long-form friendly
  • Balanced Chinese quality and 32-language coverage
  • Free tier (10K credits) to validate first

Cons

  • HD at $100/1M isn’t cheap — value comes from Turbo and clone fees
  • Streaming limited to ≤5,000-character requests
  • China-region deployments; extra compliance review for global enterprises
  • Expressiveness controls (emotion tags) behind v3/CosyVoice

Inworld Realtime TTS-2 Closed · Commercial API7.9/10

Org: Inworld AI (US; games/agents focus)
Focus: Real-time voice-agent TTS for games and interactive narrative
Arena: Controlled Voice #2 (Elo 1132); Flash tier $10.4/1M
Source: inworld.ai

The quiet champion of gaming and AI-companion scenarios: Realtime TTS-2 ranks #2 on the arena at 40% of Cartesia’s price ($20.8/1M), and the Flash tier runs $10.4/1M — “top-two quality + below-median price” is the best value in the agent track. General dubbing/audiobook ecosystems lag ElevenLabs.

Pros

  • Arena #2 — third-party quality backing
  • Cheapest among the top two (Flash $10.4/1M)
  • Deep game-engine/agent-framework integrations
  • Strong emotion and roleplay tuning

Cons

  • Agent/game focus; offline dubbing pipelines weak
  • Moderate language coverage (28)
  • Basic clone-management features
  • Lower brand awareness; fewer community resources

OpenAudio S1 / S2.1 Pro (Fish Audio) Closed API + Open Mini7.6/10

Org: Fish Audio
Focus: Commercial API ($15/1M) + open-weight S1 Mini
Arena: S2.1 Pro 1011 / S1 Mini 1010 — the open-weight tier closest to ElevenLabs Multilingual v2 (1008)
Source: fish.audio

A hybrid “half open, half closed” play: S1 Mini ships open weights (self-hostable), while commercial S2.1 Pro ties ElevenLabs’ older flagship on the arena. Solid EN/ZH/JA coverage with native emotion markers (e.g., (angry)) and multi-speaker dialogue. A natural “try open first, switch to API at scale” path.

Pros

  • Dual track: open Mini + commercial API, frictionless migration
  • Open-weight first tier on the arena (nears ElevenLabs’ old flagship)
  • Friendly $15/1M; emotion markers and multi-speaker built in
  • Clones from 10–30s

Cons

  • Open Mini weights not fully commercial-free (CC-BY-NC family)
  • Enterprise compliance tooling (consent/watermarking) behind ElevenLabs
  • Few long-tail languages
  • Small company — SLA and stability to watch

4. Head-to-Head: Open vs Closed, Dimension by Dimension

Dimension Open-Source Camp Closed Camp Edge
Similarity (top tier) Zero-shot reaches the “hard to catch in blind tests” bar, but lacks a fine-tuning tier PVC-class fine-tuned cloning remains the near-indistinguishable ceiling Closed
Naturalness / expressiveness CosyVoice emotion tags and IndexTTS disentanglement match the feature forms v3 audio tags + arena-leading quality Closed, slightly
Cloning efficiency F5-TTS 10s, Chatterbox 5–10s Cartesia 3s, MiniMax 5–10s Even (Cartesia’s extreme lead)
Cross-lingual XTTS 17, Chatterbox 23; long-tail quality varies ElevenLabs 70+, Google 380+ preset voices Closed
Real-time / streaming CosyVoice 150ms first packet (open’s only), Chatterbox streaming Cartesia 40ms, ElevenLabs Flash 75ms Closed (extremes); open is usable
Cost Self-host ≈ electricity; no metering $6.6–100/1M; MiniMax $1.50/voice lowest entry Open (at scale)
License & data sovereignty MIT/Apache: commercial, private, local — voiceprints stay in-house Data to cloud; compliance packaging (watermarks/verification) is the counterweight Open (weights), closed (tooling)
Engineering convenience Deployment, batching, monitoring all DIY API-ready, SLAs, consoles, integration ecosystems Closed

Closed wins five of eight, open wins three — but as in the digital-human report, the three open wins (cost, license, sovereignty) are precisely the most expensive items for scaled commercial deployment.

5. Scoring & Head-to-Head Evaluation

Scoring disclaimer: the scores and star ratings below are this report’s unofficial evaluation, compiled from public materials, the Artificial Analysis Controlled Voice Arena, vendor documentation, and community testing. They are not official benchmark numbers. Arena Elo uses a unified English protocol; multilingual capability is assessed from public specs and community feedback.

5.1 Criteria & Weights (customized for the voice-cloning task)

Criterion Weight What it measures
Voice similarity 25% How distinguishable the clone is from the real speaker (zero-shot and fine-tuned assessed accordingly)
Naturalness & expressiveness 20% Listening quality, emotion/style controllability
Cross-lingual 15% Language coverage and cross-language clone quality — a hard requirement for video localization
Sample efficiency 10% Reference-audio duration and onboarding cost
Speed & real-time 10% RTF, first-packet latency, streaming
Cost 10% API pricing / self-host marginal cost / clone-fee structure
License & compliance 10% Commercial license permissiveness + consent/watermark tooling

5.2 Overall Scoreboard

Model Camp Arena Elo* Score Stars Progress
Cartesia Sonic 3.6 Closed 1143 8.2 ★★★★★
CosyVoice 2 / 3 Open Not listed 8.2 ★★★★★
ElevenLabs v3 Closed 1069 8.1 ★★★★☆
Chatterbox / Turbo Open 937 7.9 ★★★★☆
Inworld Realtime TTS-2 Closed 1132 7.9 ★★★★☆
MiniMax Speech 2.8 Closed 1030 7.8 ★★★★☆
F5-TTS Open Not listed 7.8 ★★★★☆
IndexTTS-2 Open (restricted license) Not listed 7.8 ★★★★☆
OpenAudio S1 / S2.1 Pro Hybrid 1011 7.6 ★★★★☆
VibeVoice Open Not listed 7.5 ★★★★★
Higgs Audio V3 Open 963 7.5 ★★★★★
OpenVoice v2 Open 787 7.4 ★★★★★
XTTS-v2 Open (non-commercial) 839 7.3 ★★★★★
GPT-SoVITS Open Not listed 7.3 ★★★★★

* “Not listed” means no same-protocol arena data — it is not a quality ranking. The composite includes cost and licensing, so it deliberately differs from a pure-quality Elo order — e.g., ElevenLabs ranks higher on pure quality than CosyVoice, but CosyVoice’s free price and Apache license flip the composite. Differences within ±0.3 are ties.

5.3 Seven-Criterion Detail

Legend: ✓ strength · ◐ average · ✗ weakness

Model Similarity
25%
Naturalness
20%
Cross-Lingual
15%
Efficiency
10%
Speed/RT
10%
Cost
10%
License
10%
Cartesia Sonic 3.6 9.0 9.0 7.5 9.5 10 6.0 5.0
CosyVoice 2 7.5 7.5 8.0 8.5 8.5 10 9.5
ElevenLabs v3 9.5 9.5 9.5 8.0 6.0 4.0 6.0
Chatterbox 7.5 8.0 7.0 8.5 8.0 8.5 9.0
Inworld TTS-2 8.5 8.5 7.0 8.0 9.0 8.5 5.0
MiniMax Speech 2.8 8.5 8.5 8.0 8.5 7.5 8.0 4.0
F5-TTS 7.5 7.5 6.0 9.0 7.0 10 9.0
IndexTTS-2 8.5 8.5 6.5 8.5 7.0 10 4.5
OpenAudio S1/S2 8.0 8.0 7.5 8.0 6.5 9.0 5.5
VibeVoice 7.5 7.5 6.0 8.0 5.0 10 9.5
Higgs Audio V3 7.0 7.5 7.0 7.5 6.5 10 8.0
OpenVoice v2 6.5 6.0 6.0 9.0 9.0 10 9.0
XTTS-v2 6.5 6.5 8.0 8.5 8.0 10 5.0
GPT-SoVITS 8.0 7.0 5.0 6.0 6.5 10 9.0

5.4 Category Champions

Quality / Expressiveness Champion
ElevenLabs v3 + PVC

Audio-tag expressiveness + the near-indistinguishable fine-tuned ceiling.

Third-Party Arena Champion
Cartesia Sonic 3.6 (Elo 1143)

#1 on a unified protocol across all players — and also the lowest latency.

Open All-Round Champion
CosyVoice 2 (Apache-2.0)

The most feature-complete open stack: cloning + emotion + streaming.

Sample-Efficiency Champion
Cartesia (3s) / F5-TTS (10s)

Closed extreme at Cartesia; the practical open line at F5-TTS.

Emotion-Control Champion
IndexTTS-2 (disentanglement)

A’s voice with B’s emotion — a unique open trick (mind the license).

Value Champion
MiniMax ($1.50/voice)

One-time clone fee + Turbo $60/1M — the cheapest closed path at scale.

Chinese Fine-Tune Champion
GPT-SoVITS (MIT)

Fine-tune on 1 minute for top Chinese similarity (hands-on training).

Long-Form Champion
VibeVoice (90 min / 4 speakers)

Podcast/audiobook-length multi-speaker audio, MIT-licensed.

5.5 Head-to-Head Verdicts

① Multilingual Video Dubbing: ElevenLabs v3 vs CosyVoice 2

Video localization demands 30+ languages, expressiveness, and batch stability. v3’s 70+ languages, audio tags, and dubbing workbench are turnkey; CosyVoice is free but English tuning and plumbing are on you. Under ~500K chars/month, ElevenLabs saves ops; at scale, self-hosted CosyVoice wins by an order of magnitude. Verdict: ElevenLabs at low volume, CosyVoice at high volume.

② Real-Time Voice Agents: Cartesia Sonic 3.6 vs MiniMax 2.8

For agents, latency is the experience. Cartesia’s 40ms Turbo plus arena #1 and mature agent-stack integrations beat MiniMax, whose streaming caps at 5,000 characters. Winner: Cartesia Sonic 3.6 — budget alternative: Inworld Flash ($10.4/1M).

③ Local Zero-Shot Cloning: F5-TTS vs XTTS-v2

Both classics on a single RTX 4090. F5-TTS wins quality, sample efficiency (10s), and an MIT license; XTTS-v2 wins setup convenience and 17 languages, but CPML’s commercial ban is a hard flaw. Winner: F5-TTS — keep XTTS-v2 for prototypes and personal use only.

④ High-Similarity Chinese Cloning: IndexTTS-2 vs GPT-SoVITS

IndexTTS-2 is zero-shot with emotion disentanglement, usable in minutes; GPT-SoVITS reaches higher similarity after fine-tuning with 1 minute of data and ships MIT-clean. Verdict: speed and emotion → IndexTTS-2 (negotiate commercial terms first); clean licensing and peak similarity → GPT-SoVITS.

6. Scenario Recommendations

Video Translation / Multilingual Dubbing

ElevenLabs v3 or MiniMax 2.8

Pair with lip-sync pipelines for 30+ language localization; start on ElevenLabs’ ecosystem, graduate to MiniMax or self-hosted CosyVoice at scale.

Real-Time Voice Agents / Support Bots

Cartesia Sonic 3.6

40ms latency plus arena-#1 quality; budget pick: Inworld Flash ($10.4/1M).

Audiobooks / Podcast Batch Production

ElevenLabs PVC or VibeVoice

For “hardest to distinguish,” use ElevenLabs professional cloning; for budget multi-speaker long-form, VibeVoice (90 min / 4 speakers, MIT) locally.

VTubing / Game Characters

IndexTTS-2 + GPT-SoVITS

Emotion-disentangled performance plus fine-tuned similarity; resolve IndexTTS-2 licensing before commercial use, or stay all-MIT with GPT-SoVITS/RVC.

Timbre Swap on Existing Audio / AI Covers

OpenVoice v2 / RVC

The voice-conversion route: keep the performance, swap the timbre; RVC for singing, OpenVoice v2 for speech — both MIT.

Privacy-Sensitive / On-Prem

CosyVoice 2 + F5-TTS

A fully local pipeline (Apache/MIT) — voiceprints never leave the machine. The only acceptable answer for finance, healthcare, and government.

7. Decision Framework

The Five-Question Method

Q1: Re-speak or re-voice? Generating new narration → zero-shot/fine-tuned cloning (TTS route). Keeping a performance but swapping timbre → voice conversion (OpenVoice/RVC/Respeecher).

Q2: Can data leave the machine? Voiceprints are biometric data. If not → open self-hosting (CosyVoice/F5-TTS/Chatterbox). If yes → continue.

Q3: Real time? Live agents/streaming → Cartesia (40ms) or Inworld; offline content → v3/MiniMax/CosyVoice batch.

Q4: What volume? < 500K chars/month → subscriptions ($5–22/mo); millions+ → API negotiation or open self-hosting; MiniMax’s $1.50/voice is the lowest trial threshold.

Q5: Is the license clean? Before launch check three things: model license (XTTS’s CPML and IndexTTS-2’s non-commercial terms are the classic traps), the consent chain (written permission from the voice owner), and labeling duties (AI-generated content disclosure).

8. Trends

  • Architectural diversification: SSMs challenge DiTs. Cartesia’s state-space models took both quality #1 and 40ms real time, breaking the old “diffusion = slow, autoregressive = fast” rule; flow matching (F5-TTS) is the new open-source default.
  • Expressiveness is the new battlefield. ElevenLabs v3’s audio tags, CosyVoice’s emotion tags, IndexTTS-2’s timbre/emotion disentanglement, Chatterbox’s exaggeration dial — the competition moved from “sounding like” to “performing like.”
  • Open weights beat closed in blind tests for the first time. Chatterbox-Turbo’s 65.3% over ElevenLabs (vendor-run) plus open S1 Mini tying ElevenLabs’ old flagship on the arena — the open “good enough” line has crossed most commercial scenarios.
  • Prices keep breaking floors. ElevenLabs cut up to 55% in May 2026; Speechify at $6.6/1M, Inworld Flash $10.4/1M, Fish $15/1M — the per-character floor fell an order of magnitude within a year, blurring subscription vs metered boundaries.
  • From single models to pipelines. Leading products embed cloning into larger pipelines (dubbing + lip sync + translation; agents + RAG) — single-API differentiation is compressing; latency, concurrency, and compliance become the new selling points.

9. Compliance & Ethics Notes (Voice Cloning’s Required Reading)

  • Voiceprints are biometric data. Cloning someone’s voice requires explicit consent — ElevenLabs faced a BIPA class action (May 2026) over IVC’s checkbox self-attestation, which Consumer Reports called “no meaningful technical barrier.” Commercial products should use recorded-statement verification (Voice-Captcha-style).
  • Fraud is the biggest negative externality. Deepfake-driven vishing grew over 1,600% year-over-year, enterprises lose ~$680K per attack on average, and the Arup case cost $25.6M in one call. The stronger the cloning capability, the more it needs paired watermarking and detection (ElevenLabs AI Speech Classifier, Resemble’s detector).
  • Labeling duties cut both ways. EU AI Act Article 50 (effective Aug 2026) requires prominent AI-content disclosure; China’s deep-synthesis rules mandate explicit and implicit labels. Embed hidden watermarks and add AI labels to cloned-voice output.
  • Commercial license red lines. XTTS-v2 (CPML), IndexTTS-2 (non-commercial), parts of Fish Speech weights (CC-BY-NC) — open ≠ commercially free. Verify licenses and voice-data provenance model by model before launch.

10. Key Takeaways

  • “Arena #1” changes hands — ElevenLabs’ default advantage is gone. Cartesia Sonic 3.6 (1143) and Inworld TTS-2 (1132) lead ElevenLabs v3 (1069) on a unified third-party protocol. Select by testing your own material, not by brand default.
  • The open-vs-closed quality gap is now inside blind-test noise. Chatterbox’s blind-test win (vendor-run) and open S1 Mini tying ElevenLabs’ old flagship mean the closed camp’s hard advantages are down to the PVC fidelity ceiling, 70+ languages, and compliance tooling.
  • A new cost structure has appeared. MiniMax’s $1.50 one-time clone fee breaks subscription lock-in; self-hosted marginal cost approaches zero — at scale, voice is the cheapest media asset to duplicate.
  • This is the license-trap era. Non-commercial licenses (XTTS, IndexTTS-2) and compliance litigation (ElevenLabs BIPA) arrived together — the first 2026 selection question isn’t quality, it’s “can I legally sell this voice.”
  • The three roadmaps don’t substitute for each other. Localization needs voice conversion to preserve performance; new content needs zero-shot for iteration speed; IP-level voice assets need fine-tuning for the ceiling. Pick the route before comparing models.
  • Chinese teams anchor both the open and value ends. CosyVoice (Apache), F5-TTS (MIT), IndexTTS-2, GPT-SoVITS, Fish Audio, MiniMax — the world’s open voice-cloning weights and lowest price bands are supplied largely by Chinese teams.

Run Voice Cloning on Your Own Machine

DeepVideo is a locally running AI video translation and dubbing client: Voice Clone in 30+ languages, 100+ preset voices, TTS + lip sync — voiceprint data is processed on your device, never uploaded to the cloud.

Try DeepVideo Free

Realistic AI Video Dubbing at Just $0.17/min
High quality · Low price · Local client · Security

Free tier: 18 minutes total + 2 minutes daily · Local processing, no cloud upload

DeepForgeHub Research · Voice Cloning Models Global Comparison · September 2026

deepforgehub.com · Compiled from public sources; scores are unofficial evaluations

Newsletter Updates

Enter your email address below and subscribe to our newsletter

Leave a Reply

Your email address will not be published. Required fields are marked *