Speaker Diarization & Multi-Speaker Recognition
A Global Comparison of Open-Source and Commercial Models — September 2026
DER Accuracy · Speaker Capacity · Real-Time · Licensing · Cost
Executive Summary
Multi-speaker recognition — speaker diarization — went through a generational shift in 2025–2026: from the cascaded design where you bolt a separation module onto an ASR, to end-to-end speaker-attributed models that emit who spoke when and what they said in a single forward pass. The field now has three battlegrounds: accuracy (whose DER is lowest), capacity (how many speakers the model can hold), and deployment (whether audio can leave your building).
The global picture compresses to one sentence: on the open-source side, pyannote and Alibaba’s 3D-Speaker split the English and Chinese halves of the world, while NVIDIA’s Sortformer is the strongest model inside a narrow envelope of “≤4 speakers, mostly English”. On the commercial side, pyannoteAI’s Precision-2 is the accuracy ceiling across every published benchmark domain, while AssemblyAI and Deepgram sell diarization as a feature of transcription — cheap, convenient, and capped.
Key findings:
- DER numbers cannot be compared across scoring protocols: collar tolerance and whether overlapping speech is scored can swing a model by several points
- Speaker limits are hard constraints: Sortformer 4, Google Chirp 8, AWS 10, Azure 4–10, AssemblyAI 20, ElevenLabs 32; pyannote and 3D-Speaker have no ceiling
- Chinese audio has a local winner: 3D-Speaker posts 10.30% DER on AISHELL-4, beating pyannote 3.1’s 12.2%
- Microphones matter more than models: on in-the-wild video (AVA-AVD) every system lands near 50% DER
- License trap: NeMo Sortformer weights are CC-BY-NC-4.0 — not commercially usable; pyannote weights are CC-BY-4.0, usable with attribution
- Vendor benchmarks contradict each other: pyannoteAI, AssemblyAI and Speechmatics each claim first place, on metrics each chose themselves
1. First, Separate Three Tasks That Get Confused
“Multi-speaker recognition” is a vague phrase covering three different jobs. Most selection mistakes start with conflating them.
| Task | Answers | Typical output | Representative systems |
|---|---|---|---|
| Speaker Diarization | Who spoke when | Timeline with speaker labels (SPEAKER_00 / 01) | pyannote, NeMo Sortformer, 3D-Speaker, WhisperX |
| Speaker-Attributed ASR | Who said what, and when | Full transcript with speaker labels | VibeVoice-ASR, AssemblyAI, Deepgram, Speechmatics |
| Speaker ID / Verification | Is this speaker Alice? | Identity label or similarity score | CAM++, ECAPA-TDNN, pyannote voiceprint |
This report focuses on the first two. The third is usually a pre- or post-processing component — segment first, then use voiceprints to decide who it was, or to replace “SPEAKER_00” with a real name.
2. The Global Landscape
2.1 Open Source
| Option | Org | Key DER | Speaker cap | Streaming | License |
|---|---|---|---|---|---|
| pyannote.audio 4.0 · Community-1 | pyannoteAI (France) | AISHELL-4 11.7% / AMI-IHM 17.0% | Unlimited | Via commercial Live-1 | Code MIT / weights CC-BY-4.0 |
| pyannote.audio 3.1 (legacy) | pyannoteAI (France) | AISHELL-4 12.2% / AMI-IHM 18.8% | Unlimited | ✗ | MIT |
| NeMo Sortformer v2 | NVIDIA | DIHARD3 14.63% (≤4 spk) / CALLHOME-2 6.27% | 4 (hard cap) | ✓ (214× real-time) | CC-BY-NC-4.0 |
| NeMo MSDD (cascaded) | NVIDIA | DIHARD3 29.40% / CALLHOME-2 11.41% | Can exceed 4 | ✓ | CC-BY-NC-4.0 |
| 3D-Speaker (CAM++ / ERes2NetV2) | Alibaba Tongyi (China) | AISHELL-4 10.30% / AliMeeting 19.73% | No hard cap | Via FunASR | Apache-2.0 |
| FunASR (CAM++ + FSMN-VAD + spectral clustering) | Alibaba (China) | AISHELL-4 13.3% | No hard cap | ✓ WebSocket | Apache-2.0 |
| VibeVoice-ASR (7B / 8.3B) | Microsoft | MLC DER 4.28% / cpWER 11.48% | Not published | Separate streaming release | Open weights |
| WhisperX + pyannote | Community | Inherits pyannote, plus alignment noise | Unlimited | ✗ | BSD-2 |
| SpeechBrain (ECAPA-TDNN) | SpeechBrain | Only comparable with oracle VAD | Unlimited | ✗ | Apache-2.0 |
| Kaldi (x-vector) | Community | Classic baseline | — | ✗ | Apache-2.0 |
Lower DER is better. Protocol warning: these figures come from each model card’s strictest setting (no collar, overlap scored), but baselines and training data differ — compare magnitudes, not decimals.
2.2 Commercial Services
| Service | Price | Accuracy claim | Cap | Deployment |
|---|---|---|---|---|
| pyannoteAI Precision-2 | €0.112/hr (Dev) / €0.096 (Starter); €19/mo incl. 170 hr | Lowest DER in all ten benchmark domains | Unlimited | Cloud + on-prem |
| pyannoteAI Live-1 | Priced separately | Streaming scenarios | Unlimited | Cloud |
| AssemblyAI Universal-3.5 Pro | $0.21/hr + diarization $0.02/hr | cpWER 30.17 (own benchmark) | 20 async / 10 streaming | Cloud |
| Deepgram Nova-3 | $0.0043/min batch (≈$0.258/hr) + diarization $0.0020/min | cpWER 37.92 EN (third-party figure) | No hard cap | Cloud + self-host |
| Speechmatics Ursa 2 | ≈$1.20/hr | Strong on hard audio, multilingual | 2–20 configurable | Cloud + on-prem |
| ElevenLabs Scribe v2 | $0.22/hr batch / $0.39/hr realtime | cpWER 35.26 (third-party figure) | 32 | Cloud |
| Gladia | $0.61/hr (down to $0.20/hr on Growth) | Backed by pyannoteAI Precision-2 | — | Cloud (EU/US) |
| Rev AI | $0.02/min (human review $1.50/min) | DER 10–13% AMI | — | Cloud |
| OpenAI gpt-4o-transcribe-diarize | $0.36/hr | No published DER | Not published | Cloud |
| Google Cloud Chirp 3 | $0.016/min (dynamic batch $0.004/min) | Trails on hard audio | 8 | Cloud (V2 API) |
| Azure AI Speech | Realtime $0.0167/min, batch $0.006/min | Trails Speechmatics | 4–10 | Cloud + container |
| AWS Transcribe | $0.024/min | Diarization scored 8.3/10 | 10 | Cloud |
| Tencent Cloud (China) | Role separation ¥0.85/hr; 1:N voiceprint ¥4.2/1k calls | No published DER | — | Cloud |
| iFlytek Tingjian (China) | Free 2 hr/mo; ¥29/hr or ¥199/mo | 98%+ Chinese recognition | — | Cloud + offline |
List prices as of September 2026, all billed by audio duration. For most vendors diarization is a paid add-on, not a base feature (Deepgram +$0.0020/min, AssemblyAI +$0.02/hr).
3. Deep Dive, One by One
pyannote.audio 4.0 · Community-1 Best open source Score 8.9
The de facto standard in open source. Community-1 replaces the older segmentation and embedding models and adds VBx clustering, which substantially improves speaker assignment and speaker counting — exactly the two metrics that hurt 3.1. It also adds an “exclusive” mode that emits only the single most likely speaker at any moment, purpose-built for aligning with Whisper-style word timestamps and eliminating the most annoying failure mode in cascaded pipelines. The library sees roughly 45 million downloads a month on Hugging Face; most commercial transcription products quietly run it underneath.
Strengths
- Highest accuracy in open source: AISHELL-4 11.7% DER
- Clustering architecture, no hard speaker ceiling
- Language-agnostic — no special handling for Chinese or others
- Exclusive mode solves ASR timestamp reconciliation
- Self-host, hosted, or upgrade to Precision-2 with one parameter
Weaknesses
- Needs a GPU and MLOps — self-hosting cost is operational
- Hugging Face weights are gated; accept terms per model
- Open pipeline covers batch only; streaming means commercial Live-1
- Speaker counting still errs above 8 speakers
pyannoteAI · Precision-2 / Live-1 Accuracy ceiling Score 9.1
On every publicly checkable benchmark, Precision-2 takes the lowest DER — including the vendor’s own ten-domain DIHARD benchmark (259 recordings, ~67 hours) and an independent academic study over 196.6 hours of multilingual audio (English, Mandarin, German, Japanese, Spanish) that measured 11.2% DER, best in field. Official figures put it ~28% more accurate than the open Community-1, and 2.2–2.6× faster when self-hosted. Its distinguishing feature is voiceprints: enrol once, then turn “SPEAKER_00” into a real name and recognise the same person across files. For multi-speaker video dubbing — where each character’s voice must map to a fixed speaker — that capability is worth a lot.
Strengths
- Lowest DER in all ten domains — no weak spot
- No speaker cap, overlap detection included
- Voiceprints enable cross-file identity recognition
- STT orchestration returns an attributed transcript in one call
- On-prem option; customer audio never used for training
Weaknesses
- Pricier than AssemblyAI / Deepgram by one tier
- Benchmark is vendor-maintained; independent replication still limited
- Streaming requires Live-1; no open-source equivalent
- You still need to pair it with an ASR (or use its orchestration)
NVIDIA NeMo · Sortformer v2 King of a narrow envelope Score 7.1
Sortformer sidesteps the permutation problem by sorting speakers in arrival order, emitting labels end-to-end without clustering. Inside its envelope the numbers are hard: DIHARD-III (≤4 speakers) 14.63% versus 29.40% for the cascaded MSDD, CALLHOME-2 6.27% versus 11.41% — and it is fast, with the streaming version hitting 214× real-time. There is exactly one problem, and it is fatal: the model detects a maximum of four speakers. That is architectural, not a tuning default — the output layer is built for four. Performance degrades beyond that, and a 14-speaker recording is simply out of scope.
Strengths
- Doubles MSDD accuracy inside the ≤4-speaker envelope
- End-to-end single model — no clustering pipeline to assemble
- Native overlapped-speech detection
- Streaming version at 214× real-time, very low latency
- Tight Riva SDK integration for enterprise stacks
Weaknesses
- Four speakers maximum — breaks on meetings, panels, call transfers
- Trained primarily on English; mediocre on Chinese
- CC-BY-NC-4.0 forbids commercial use
- NeMo is a research framework; Hydra configs are a learning curve
Alibaba 3D-Speaker + FunASR Best for Chinese Score 8.4
The strongest open-source answer for Chinese audio. 3D-Speaker reaches 10.30% DER on AISHELL-4 (Chinese meetings) — better than pyannote 3.1’s 12.2% and Community-1’s 11.7%; CAM++ hits 0.65% EER on VoxCeleb1-O with only 7.2M parameters, which is remarkable value. It is also the only public toolkit supporting true multimodality (audio + visual + semantic), letting mouth shape and on-screen presence assist who-is-speaking decisions. FunASR packages CAM++, FSMN-VAD and spectral clustering into one pipeline — the highest level of integration for Chinese production deployments — with a Docker image that serves streaming speaker labels over WebSocket.
Strengths
- Leads on Chinese audio: AISHELL-4 10.30% DER
- Apache-2.0 — cleanest commercial terms of any option here
- Spectral clustering, no hard speaker cap
- Streaming speaker labels supported
- Multimodal extension (lip / visual cues)
Weaknesses
- Dependency conflicts are common (fastcluster / hdbscan / libnvrtc)
- Weaker than pyannote on English and multilingual audio
- Documentation is Chinese-first; thin overseas community support
- Unstable on English benchmarks like VoxConverse
Microsoft VibeVoice-ASR / Streaming The unified paradigm Score 8.3
The most important architectural signal of 2026. The offline version ingests up to 60 minutes of audio without chunking and emits structured JSON (speaker ID + start/end + text), which removes the “speaker drift” that chunked pipelines introduce at every boundary; on the MLC-Challenge multi-speaker benchmark it posts DER 4.28% and cpWER 11.48%. The streaming version goes further — among the first LLM-based end-to-end streaming speaker-attributed ASR systems, emitting “who said what” as audio arrives, with an expected speaker-attribution latency of 2.00 seconds versus a measured 8.21 s for Azure ConversationTranscriber and 9.12 s for Google Cloud STT. It takes best or tied-best in 12 of 13 speaker-attribution settings.
Strengths
- Unified model removes cascaded alignment error
- 2.0 s attribution latency — 4× faster than cloud vendors
- 60-minute single pass, best long-form consistency
- 50+ languages, native Chinese-English switching
- Custom hot-words; deployable via vLLM
Weaknesses
- 7B/8.3B needs 17 GB+ VRAM — high barrier
- General ASR WER 7.77% — good, not class-leading
- Speaker-count ceiling not publicly documented
- Streaming loses 5–6.7 cpWER points versus offline
WhisperX + pyannote Most practical combo Score 8.3
The engineering wrapper that stitches Whisper, forced alignment and pyannote into a single command — the default answer to “I want it running today”. Its value is word-level speaker labels: not just “this segment is speaker 00”, but every word mapped back to a speaker, which subtitles and line-by-line dubbing require. The cost is two layers of error stacking: alignment adds noise, and on real messy audio (Scribie’s production set) it measures 26.31% DER — not meaningfully better than AssemblyAI’s 26.61%.
Strengths
- One command to run; the deepest pool of tutorials and answers
- Word-level timestamps + speaker labels, ideal for subtitles
- Whisper covers 99 languages
- Completely free, offline, air-gap capable
Weaknesses
- ASR and diarization errors compound
- Real messy audio can reach 26%+ DER
- Batch only, no streaming
- Whisper itself struggles on overlapped multi-speaker audio
AssemblyAI · Universal-3.5 Pro Least integration effort Score 8.6
The best developer experience in the category: one key, no tuning, and you get a speaker-labelled transcript plus summarisation and entity detection. Universal-3.5 Pro, released in July 2026, optimises for cpWER (which ties speaker labels to the actual transcribed words) and scores 30.17 in the vendor’s own comparison, ahead of ElevenLabs Scribe v2 (35.26), Gladia (36.87) and Deepgram Nova-3 EN (37.92). Speaker Identification can replace “Speaker A” with role names like “Agent” or “Customer” — but note that every one of those numbers comes from the vendor’s own test set.
Strengths
- Lowest integration cost — one API key
- 20 speakers async / 10 streaming is plenty for most products
- Role-name labels instead of generic speaker IDs
- Realtime and batch on the same interface
- Native code-switching across 18 languages
Weaknesses
- cpWER 30.17 is a vendor figure with no third-party replication
- Cloud only — no self-hosted or on-prem option
- Diarization is an ASR feature; nothing to tune when labels are wrong
- Billed by session duration, so short audio costs disproportionately more
Deepgram · Nova-3 / Flux Speed and cost leader Score 8.7
If what you want is “cheap and fast”, this is essentially the only answer. Nova-3 emits words and speaker labels from a single forward pass, is language-agnostic, has no hard speaker cap, and has the lowest batch unit price among major APIs. Flux, shipped in late 2025, goes further by folding turn detection (whose turn is it next) into the recogniser itself, removing the need for an external VAD — a structural advantage for voice agents. The trade-off is that diarization is a feature of the ASR engine, so on hard audio (crosstalk, overlap, heavy accents) accuracy visibly falls behind.
Strengths
- Lowest unit price: $0.0043/min batch
- Sub-300 ms latency, best streaming experience
- No hard speaker ceiling
- Flux builds in turn detection — voice-agent friendly
- Self-hosting available
Weaknesses
- cpWER 37.92 — last among the three-way comparison
- Accuracy drops noticeably on hard multi-speaker audio
- 36+ languages, fewer than the hyperscalers
- Diarization and several enrichments are billed separately
Speechmatics · Ursa 2 Compliance and sovereignty Score 8.2
The answer for hard audio and compliance requirements. It has the strongest reputation on accented English and on domains where a mislabelled speaker costs money — legal, medical, broadcast — and is one of the few commercial services offering on-prem and air-gapped deployment, meaning audio never leaves the internal network. In several regulated industries that single line decides procurement. The price is among the highest of the mainstream APIs (≈$1.20/hr, about 4.6× Deepgram) and it sells through enterprise channels, which rules out casual pilots.
Strengths
- Best reputation on hard audio (accents, crosstalk, multilingual)
- On-prem and air-gapped deployment
- Configurable 2–20 speakers
- Enterprise SLAs and compliance tooling
Weaknesses
- Most expensive mainstream option — ~4.6× Deepgram
- Enterprise sales only; high barrier to trial
- All accuracy rankings are vendor self-reported
- Smaller ecosystem than the hyperscalers
The Cloud Trio · Azure / AWS / Google Already on your bill
These three are valuable not for being best but for already being on your invoice. If your audio sits in S3, Blob Storage or GCS, they save you a data hop. But all three treat diarization as an accessory of a general speech platform: Google Chirp 3 caps at 8 speakers, AWS at 10, and Azure’s Fast Transcription recommends 4 (its multi-device conversation mode reaches roughly 10). There is also a commonly missed shortcut — if the recording is dual-channel (agent and customer on separate tracks), use channel identification rather than diarization: the channel already tells you who is who, no inference required, and it is free of the usual error modes.
Strengths
- Zero friction with existing cloud infrastructure
- Widest language coverage (Azure 140+, Google 125+)
- AWS and Azure offer HIPAA BAA paths
- Batch pricing as low as $0.004/min
Weaknesses
- Low speaker caps (Google 8, AWS 10)
- Accuracy trails dedicated vendors on hard multi-speaker audio
- Almost nothing tunable in the diarization path
- Vendor lock-in raises migration cost
4. Head-to-Head Matrix
| Option | Architecture | Speaker cap | Streaming | Languages | On-prem | Commercial license |
|---|---|---|---|---|---|---|
| pyannote Community-1 | Cascaded + VBx | Unlimited | ✗ | Agnostic | ✓ | CC-BY-4.0 |
| pyannoteAI Precision-2 | Cascaded + voiceprints | Unlimited | ✓ (Live-1) | Agnostic | ✓ (Enterprise) | Commercial |
| NeMo Sortformer v2 | End-to-end | 4 | ✓ | English-first | ✓ | Non-commercial |
| 3D-Speaker / FunASR | Cascaded + spectral | No hard cap | ✓ (FunASR) | Chinese-first | ✓ | Apache-2.0 |
| VibeVoice-ASR | End-to-end LLM | Not published | ✓ (streaming release) | 50+ | ✓ | Open weights |
| WhisperX + pyannote | Cascaded + alignment | Unlimited | ✗ | 99 | ✓ | BSD-2 |
| AssemblyAI U-3.5 Pro | Unified ASR | 20 / 10 | ✓ | 18 (99 via U-2) | ✗ | Commercial |
| Deepgram Nova-3 | Unified ASR | No hard cap | ✓ | 36+ | ✓ | Commercial |
| Speechmatics Ursa 2 | Unified ASR | 2–20 configurable | ✓ | 50+ | ✓ | Commercial |
| Azure / AWS / Google | Unified ASR | 4–10 | ✓ | 100+ | Azure container only | Commercial |
5. Measured Data: The Part Most Write-ups Soften
This section does one thing: puts numbers back inside the protocol they came from. Diarization is one of the few fields where a single model can be reported with a dozen different DER values, so a comparison without a protocol note means nothing.
5.1 The Power of Protocol: One Model, Three Sets of Numbers
| Model | AISHELL-4 | AMI-IHM | CALLHOME | VoxConverse | DIHARD-III |
|---|---|---|---|---|---|
| pyannote 3.1 (legacy) | 12.2% | 18.8% | 28.5% | 11.2% | 21.4% |
| pyannote Community-1 | 11.7% | 17.0% | 26.7% | 11.2% | 20.2% |
| pyannoteAI Precision-2 | 11.4% | 12.9% | 16.6% | 8.5% | 14.7% |
| 3D-Speaker | 10.30% | 21.76% (SDM) | — | 11.75% | — |
| NeMo Sortformer v2 | ~13% | ~13% | 6.27% (2 spk) | — | 14.63% (≤4 spk) |
Protocol: pyannote figures use the strictest setting — no collar, overlap scored. Two things stand out. (1) Sortformer’s 6.27% on two-speaker CALLHOME is the best number in the table, but the same model caps at four speakers on DIHARD-III — its strength is “few speakers, clean”. (2) 3D-Speaker beats every pyannote version on Chinese meetings (10.30%) yet falls to 21.76% on English AMI-SDM. Language specialisation cuts both ways.
5.2 Your Room Sets the Ceiling: Microphones Beat Models
| Audio type | Benchmark | pyannote 3.1 DER | Note |
|---|---|---|---|
| Broadcast (close-mic) | REPERE | 7.8% | Studio-grade audio; every system looks good |
| Web video | VoxConverse | 11.3% | Political debate; overlap present, quality OK |
| Chinese meetings | AISHELL-4 | 12.2% | Multi-party, far-field capture |
| Close-mic meetings | AMI (IHM) | 18.8% | One mic per person, but frequent natural turns |
| Mixed hard domains | DIHARD-III | 21.7% | Clinical, courtroom, children’s speech |
| Far-field meetings | AliMeeting | 24.4% | Single room mic, heavy reverb and crosstalk |
| In-the-wild video | AVA-AVD | 50.0% | Music, crowds, heavy overlap — half the time mislabelled |
From pyannote’s official model card (strictest setting). The conclusion is harder than any buying advice: your diarization quality is set first by microphones and acoustics, and only then by the model. The same model scores 7.8% on broadcast audio and 50% on in-the-wild video. If the budget is finite, buy one close mic per speaker before you buy a better model.
5.3 The Cost and Benefit of Unification: Attribution Latency
| System | Speaker-attribution latency | Five-set average error | Type |
|---|---|---|---|
| VibeVoice-ASR-Streaming | 2.00 s | 24.66 | End-to-end LLM, streaming |
| Gemini 3.5 Transcribe Live | — | 25.23 | Cloud streaming |
| Azure ConversationTranscriber | 8.21 s | — | Cloud streaming |
| Google Cloud STT | 9.12 s | — | Cloud streaming |
| ElevenLabs Scribe v2 Realtime | — | 41.39 | Cloud streaming |
Conditions: Microsoft VibeVoice-ASR-Streaming technical report (arXiv:2609.02812), 22-frame chunks with 4-frame lookahead; evaluation audio truncated to 8 minutes across 10 languages. The key difference is not accuracy but how soon the system dares to commit: cloud vendors revise speaker labels tens of seconds later, while the end-to-end model settles within 2 seconds. For live captioning and voice agents, that gap matters more than a point of DER.
5.4 Real Messy Audio: How Far Lab Numbers Fall
| System | N-WER (word errors) | DER (speaker errors) |
|---|---|---|
| WhisperX (large-v3 + pyannote) | 12.81% | 26.31% |
| AssemblyAI | 15.13% | 26.61% |
| Deepgram | 15.62% | 27.23% |
Conditions: Scribie production set — 26 real transcription jobs including legal depositions, multi-speaker interviews, compressed Zoom audio, phone recordings and terminology-dense content. This is the pair of numbers to remember: models that score 11% DER on podcast-clean audio commonly land at 26–27% on real call and deposition audio. Note also that DER is universally worse than word error rate — attributing speech to a speaker is harder than hearing the words — and that all three systems effectively tie here, which says the ceiling belongs to the task, not to the product.
6. Scoring and Head-to-Head Verdicts
The previous sections analysed each option qualitatively; this one puts numbers on it. This scoring is not an official benchmark — it is a composite judgment drawn from public model cards, vendor documentation, independent academic evaluation and production measurements. Seven dimensions are each scored out of 10 and weighted into a single overall score. Results will vary with audio conditions and deployment model, so treat this as a starting point for ranking, not the final word.
6.1 Scoring Dimensions and Weights
| Dimension | Weight | What it measures |
|---|---|---|
| Accuracy | 30% | DER / cpWER on public benchmarks, plus stability on real messy audio |
| Speaker capacity | 10% | Whether the speaker ceiling is hard or tunable; failure at high speaker counts |
| Language coverage | 15% | Language-agnostic design, Chinese performance, code-switching |
| Real-time capability | 10% | Streaming support, attribution latency, whether an external VAD is needed |
| Usability | 10% | Integration cost, dependency complexity, docs, whether you must wire up your own ASR |
| Cost | 15% | Effective per-hour price, or the GPU and engineering cost of self-hosting |
| Licensing & compliance | 10% | Clarity of commercial terms, on-prem availability, training-data commitments |
6.2 Overall Scoreboard
| Rank | Option | Overall score | Stars | One-line verdict |
|---|---|---|---|---|
| 1 | pyannoteAI Precision-2 commercial |
9.1
|
★★★★★ | Wins all ten benchmark domains; no capacity or compliance gap — you just pay |
| 2 | pyannote Community-1 open source |
8.9
|
★★★★★ | Best open source, upgradeable to the commercial model — you pay in ops |
| 3 | Deepgram Nova-3 commercial |
8.7
|
★★★★★ | First on speed and first on price; accuracy is what it trades away |
| 4 | AssemblyAI U-3.5 Pro commercial |
8.6
|
★★★★★ | Easiest to integrate, best unified transcript UX; accuracy awaits third-party proof |
| 5 | Alibaba 3D-Speaker / FunASR open source · China |
8.4
|
★★★★★ | Best for Chinese and fully Apache-2.0; environment setup is the barrier |
| 6 | WhisperX + pyannote open-source pipeline |
8.3
|
★★★★★ | Practical one-command combo; you pay in compounded error |
| 7 | VibeVoice-ASR open source · Microsoft |
8.3
|
★★★★★ | Most advanced architecture, leading long-form and streaming attribution; heavy VRAM cost |
| 8 | Speechmatics Ursa 2 commercial |
8.2
|
★★★★★ | The compliance and hard-audio answer; price is the barrier |
| 9 | NeMo Sortformer v2 open source · NVIDIA |
7.1
|
★★★★★ | Superb at ≤4-speaker English, but the 4-speaker cap and non-commercial licence hold it down |
Overall score = sum of dimension scores × weights (out of 10). Ranks 6 and 7 tie at 8.3 — WhisperX wins on usability and ecosystem, VibeVoice on architecture and long-form consistency.
6.3 Seven-Dimension Breakdown
| Option | Accuracy | Capacity | Language | Real-time | Usability | Cost | Compliance | Overall |
|---|---|---|---|---|---|---|---|---|
| Precision-2 | 9.8 | 10.0 | 9.5 | 9.0 | 8.5 | 7.0 | 9.0 | 9.1 |
| Community-1 | 9.3 | 10.0 | 9.5 | 6.5 | 6.5 | 9.5 | 9.5 | 8.9 |
| Deepgram Nova-3 | 8.0 | 10.0 | 8.0 | 10.0 | 9.5 | 9.5 | 7.5 | 8.7 |
| AssemblyAI U-3.5 | 8.5 | 8.5 | 8.5 | 8.5 | 10.0 | 8.5 | 8.0 | 8.6 |
| 3D-Speaker / FunASR | 8.5 | 9.5 | 7.5 | 7.5 | 6.0 | 9.5 | 10.0 | 8.4 |
| WhisperX + pyannote | 7.8 | 10.0 | 9.5 | 5.0 | 7.0 | 9.5 | 9.0 | 8.3 |
| VibeVoice-ASR | 8.5 | 8.0 | 8.5 | 9.0 | 7.5 | 8.0 | 8.5 | 8.3 |
| Speechmatics Ursa 2 | 8.5 | 8.0 | 9.0 | 8.0 | 8.0 | 6.0 | 9.5 | 8.2 |
| NeMo Sortformer v2 | 8.5 | 3.0 | 6.0 | 9.5 | 6.0 | 9.0 | 4.0 | 7.1 |
Green marks the top tier in each column, amber a clear weakness, red the weakest. Three signals. (1) Sortformer’s accuracy (8.5) and capacity (3.0) differ by 5.5 points — it is not a weak model, it is a strong model locked inside an envelope. (2) Every open-source option loses in the “real-time” column: streaming has long been a commercial-side capability only. (3) Compliance separates Sortformer on its own — a non-commercial licence is a veto in commercial settings.
6.4 Category Champions
Lowest DER in all ten DIHARD domains; 11.2% best-in-field in an independent academic study across 196.6 hours and five languages.
AISHELL-4 11.7%, AMI-IHM 17.0% — ahead of most commercial services on public benchmarks.
10.30% DER on AISHELL-4, better than any pyannote release; CAM++ reaches 0.65% EER with 7.2M parameters.
Clustering architectures have no hard speaker ceiling; Sortformer’s 4, Google’s 8 and AWS’s 10 are architectural limits.
Deepgram streams under 300 ms; VibeVoice attributes speakers in 2.0 s — four times faster than Azure.
$0.0043/min is the commercial floor; self-hosting Community-1 costs one modern NVIDIA GPU.
Apache-2.0 with no extra terms; community options sit between “attribute us” and “no commercial use”.
On-prem and air-gapped options keep audio inside the network; Azure containers are the runner-up.
6.5 Head-to-Head Verdicts: Four Matchups
① Accuracy: three protocols, three winners
pyannoteAI claims all ten domains on DER; AssemblyAI claims 30.17 cpWER against Deepgram’s 37.92; Speechmatics claims 25% ahead of its nearest competitor. All three are true and all three are partial — each picked the metric and test set that favours it. The only robust conclusion: dedicated diarization models (the pyannote family) reliably beat diarization-as-an-ASR-feature on hard audio.
② Open source: three constraint lines
Community-1 wins on accuracy + no speaker ceiling; 3D-Speaker wins on Chinese + Apache-2.0; Sortformer wins on speed (214× real-time) but is locked by the 4-speaker cap and a non-commercial licence. Ask three questions first — how many speakers, which language, commercial or not — and two of the three are eliminated.
③ Architecture: unified vs cascaded
Cascaded (WhisperX + pyannote) wins on flexibility, swappable parts and word-level timestamps, but compounds two layers of error: 26.31% DER on real messy audio. Unified (VibeVoice / AssemblyAI) wins on eliminating alignment error and attributing in 2.0 s versus 8–9 s, at the cost of a black box you cannot tune when it errs.
④ Cost: the self-host breakeven
By published figures, self-hosting only beats API pricing above roughly 14,000 audio hours per month (including GPU and engineering). Below that, hosted APIs are cheaper — but if audio cannot leave your network, self-hosting stops being a choice and becomes the only option. pyannoteAI’s €0.035/hr hosted open-source tier is the middle path: no GPU, no model lock-in.
Verdict summary: there is no all-round champion, only a best fit per scenario —
- Accuracy first, budget available, on-prem needed: pyannoteAI Precision-2 (9.1)
- Open source preferred, ops accepted: pyannote Community-1 (8.9)
- Speed and lowest unit price: Deepgram Nova-3 (8.7)
- Zero ops, fastest to ship: AssemblyAI Universal-3.5 Pro (8.6)
- Chinese audio with Apache-2.0 clarity: Alibaba 3D-Speaker / FunASR (8.4)
- Word-level labels for subtitles and dubbing: WhisperX + pyannote (8.3)
- Long-form single-pass and streaming attribution: VibeVoice-ASR (8.3)
- Hard compliance and sovereignty: Speechmatics (8.2)
- English ≤4 speakers, efficiency only (non-commercial): NeMo Sortformer v2 (7.1)
7. Engineering Practice: Six Steps That Decide the Outcome
The model is one link in the chain. Miss any of the following and an 11% DER becomes 30%.
| Step | What it decides | Practical guidance |
|---|---|---|
| Microphones and capture | The accuracy ceiling | One close mic per speaker beats a single room mic. The same model scores 7.8% close-mic and 50% in the wild — no model recovers lost room acoustics |
| VAD preprocessing | False alarms | Run FSMN-VAD or MarbleNet first to strip music, noise and silence; false alarms fall markedly |
| Overlap handling | Success on crosstalk | Ranking: Sortformer (native) > pyannote (overlap-aware) > pure clustering. For meetings, favour models that handle overlap |
| Speaker-count priors | Counting errors | Always pass num_speakers when known; when unknown, give min/max bounds. Both beat letting the model guess |
| Voiceprint enrolment | IDs to real names | Build a voiceprint store with CAM++ or ECAPA-TDNN to map SPEAKER_00 to a real identity; pyannoteAI voiceprints are the managed route |
| ASR reconciliation | Word-level label quality | Use pyannote’s exclusive mode so only one speaker is active per moment, preventing overlap regions from misattributing words |
The other half of the cost story: the self-hosting bill is not just GPU. One modern NVIDIA card runs Community-1, but you also own model version updates, expired Hugging Face tokens (which break the whole pipeline silently), dependency conflicts (3D-Speaker has long been stuck on fastcluster / libnvrtc), and batch scheduling. Hosted APIs are selling those, not just the model.
A shortcut people miss: if the recording is dual-channel (agent and customer on separate tracks), use channel identification instead of diarization. The channel tells you who is who with zero inference and zero error rate. AWS and Azure both build this in; many teams pay for diarization they never needed.
8. Recommendations by Scenario
Meeting transcription, unknown headcount
No speaker ceiling, language-agnostic — holds from 5 to 15 participants. One GPU, or hosted at €0.035/hr.
Call centre QA and conversation analytics
Dual-channel calls need channel identification, not diarization. For single-channel crosstalk precision, go Precision-2.
Chinese meetings and podcasts
10.30% DER on AISHELL-4 is the best in the field, Apache-2.0 for commercial use, streaming labels included.
Live captioning and voice agents
The former builds in turn detection under 300 ms; the latter attributes speakers in 2.0 s, four times faster than cloud vendors.
Video dubbing and multi-character translation
Voiceprints bind characters to fixed voices, then the diarization output feeds TTS for multi-character dubbing.
Subtitles needing word-level labels
The only out-of-the-box combination that yields word timestamps plus speakers — SRT straight out.
Long-form single-pass (60 minutes)
No chunking, no stitching, no speaker drift; structured JSON with timestamps out.
Audio cannot leave the building
Both support air-gapped deployment; in China, self-hosted 3D-Speaker is the equivalent route.
Tightest budget, just get started
Fully open source with zero API cost on a consumer GPU; validate on a small sample before scaling.
Decision Framework
1. Count speakers first: ≤4 and English → Sortformer works (mind the licence); 5–20 → 3D-Speaker or pyannote; 20+ → clustering only (pyannote, Deepgram)
2. Then language: Chinese-first → 3D-Speaker / FunASR; multilingual → the pyannote family (agnostic); English-only → widest choice
3. Can audio leave your network? No → self-host or Speechmatics / pyannoteAI Enterprise; Yes → hosted APIs save the ops
4. Commercial use? Yes → rule out Sortformer (non-commercial), honour CC-BY-4.0 attribution, or pick Apache-2.0 for the cleanest terms
5. Need real-time? Yes → Deepgram (fastest) or VibeVoice-Streaming (fastest attribution); No → batch saves 3–6×
6. Is the recording dual-channel? Yes → use channel identification first and skip the diarization bill
7. Need word-level labels? Yes → WhisperX + pyannote, or pyannote’s exclusive mode
8. Still undecided? → go by the overall scores: Precision-2 (9.1) for accuracy, Community-1 (8.9) for open source, Deepgram (8.7) for speed and cost
9. Pitfalls and Compliance
1. Always quote DER with its protocol. An 11% with no collar and overlap scored is not the same number as an 11% with a collar and overlap ignored. Look for those two switches before reading any comparison table — vendors routinely report the flattering side.
2. Establish the speaker ceiling early. Sortformer’s 4, Google’s 8, AWS’s 10 and Azure’s 4–10 are architectural limits, not tuning issues. One extra participant invalidates the whole approach — eliminate this risk before the POC, not after.
3. A non-commercial licence is a veto. NeMo Sortformer weights are CC-BY-NC-4.0 and cannot be used commercially. Teams that ship it into production and discover this at legal review face very expensive rework. For commercial projects choose Apache-2.0 (3D-Speaker / FunASR) or an explicit commercial licence.
4. CC-BY-4.0 requires attribution. pyannote Community-1 weights are commercially usable with an attribution notice in your product. Code under MIT and weights under CC-BY are two separate licences — do not read only the code half.
5. Read the training-data clause. pyannoteAI states customer audio is never used for training and publishes retention policy per tier; most vendors make no such commitment. For sensitive audio, put this in the contract.
6. Never trust a single vendor benchmark. Three vendors each claiming first place is normal market behaviour. The correct move: take 20–50 hours of your own real audio (including the worst of it), score every candidate with the same pyannote.metrics protocol, then decide.
7. Do not neglect VAD and silence. False alarms — noise counted as speech — are one of the three components of DER. A clean VAD in front often helps more than a model upgrade.
8. Voiceprints trigger biometric-privacy law. Using voiceprints to identify people brings biometric data rules into play: sensitive personal information with separate-notice consent obligations in China, special-category data under GDPR in the EU, and class-action exposure under BIPA-style US state laws. Clear it with legal before shipping identity features.
Key Findings
1. The field is moving from a separation module to a unified capability. The 2026 dividing line is speaker-attributed ASR — VibeVoice compresses ASR, diarization and timestamps into one forward pass; AssemblyAI and Deepgram build labels into the recogniser. The cascaded pipeline’s alignment error and latency disadvantage will only grow, though its swappable parts remain the right answer for some products.
2. The accuracy ceiling belongs to the pyannote family, and open source sits one parameter away from commercial. Community-1 leads open source, Precision-2 leads every benchmark domain, and both share one integration — a rare “validate free, upgrade seamlessly” path.
3. Chinese audio has its own winner. 3D-Speaker posts 10.30% DER on AISHELL-4, better than any pyannote release, under a fully commercial Apache-2.0 licence. Language specialisation cuts both ways — it drops to 21.76% on English AMI-SDM.
4. The speaker ceiling is the first hard filter. Four, eight and ten are not performance differences; they are the line between usable and unusable. Count your speakers before you read the accuracy table.
5. Microphones decide more than models do. The same model scores 7.8% on broadcast audio and 50% in the wild. When budget is finite, spend it on capture first.
6. Real audio is far harsher than any leaderboard. Production measurements cluster around 26–27% DER, and speaker errors are consistently worse than word errors. Any POC must run on your own messy data; clean-sample numbers carry no information.
7. Licensing is the buried landmine. Sortformer’s non-commercial terms, CC-BY attribution duties and biometric-privacy obligations around voiceprints can each stop a technically successful project before launch. Put them on the same table as accuracy when you decide.
Put Multi-Speaker Audio to Work
Diarization is the easy half; the hard half is keeping every voice intact across languages. DeepVideo by DeepForgeHub translates any video into 30+ languages with Voice Clone and lip-sync — a Windows & Mac desktop app that runs locally, so your footage never leaves your machine.
Free tier: 18 minutes total + 2 minutes daily · Local processing, no cloud upload
© 2026 DeepForgeHub Research. Sources: pyannoteAI public benchmarks and model cards, NVIDIA NeMo model cards and technical reports, Alibaba 3D-Speaker / FunASR documentation, Microsoft’s VibeVoice-ASR technical report (arXiv:2609.02812), official documentation and pricing pages from AssemblyAI, Deepgram, Speechmatics, ElevenLabs, Gladia and Rev AI, AWS / Azure / Google Cloud documentation, and third-party production benchmarks (Scribie, Neosophie). Prices are list prices; DER protocols are noted inline. As of September 2026.

