Speaker Diarization Models Report 2026: 10 Solutions Compared

Speaker Diarization Models Research Report 2026

Speaker Diarization & Multi-Speaker Recognition

A Global Comparison of Open-Source and Commercial Models — September 2026

DER Accuracy · Speaker Capacity · Real-Time · Licensing · Cost

Executive Summary

Multi-speaker recognition — speaker diarization — went through a generational shift in 2025–2026: from the cascaded design where you bolt a separation module onto an ASR, to end-to-end speaker-attributed models that emit who spoke when and what they said in a single forward pass. The field now has three battlegrounds: accuracy (whose DER is lowest), capacity (how many speakers the model can hold), and deployment (whether audio can leave your building).

The global picture compresses to one sentence: on the open-source side, pyannote and Alibaba’s 3D-Speaker split the English and Chinese halves of the world, while NVIDIA’s Sortformer is the strongest model inside a narrow envelope of “≤4 speakers, mostly English”. On the commercial side, pyannoteAI’s Precision-2 is the accuracy ceiling across every published benchmark domain, while AssemblyAI and Deepgram sell diarization as a feature of transcription — cheap, convenient, and capped.

Key findings:

  • DER numbers cannot be compared across scoring protocols: collar tolerance and whether overlapping speech is scored can swing a model by several points
  • Speaker limits are hard constraints: Sortformer 4, Google Chirp 8, AWS 10, Azure 4–10, AssemblyAI 20, ElevenLabs 32; pyannote and 3D-Speaker have no ceiling
  • Chinese audio has a local winner: 3D-Speaker posts 10.30% DER on AISHELL-4, beating pyannote 3.1’s 12.2%
  • Microphones matter more than models: on in-the-wild video (AVA-AVD) every system lands near 50% DER
  • License trap: NeMo Sortformer weights are CC-BY-NC-4.0 — not commercially usable; pyannote weights are CC-BY-4.0, usable with attribution
  • Vendor benchmarks contradict each other: pyannoteAI, AssemblyAI and Speechmatics each claim first place, on metrics each chose themselves

1. First, Separate Three Tasks That Get Confused

“Multi-speaker recognition” is a vague phrase covering three different jobs. Most selection mistakes start with conflating them.

Task Answers Typical output Representative systems
Speaker Diarization Who spoke when Timeline with speaker labels (SPEAKER_00 / 01) pyannote, NeMo Sortformer, 3D-Speaker, WhisperX
Speaker-Attributed ASR Who said what, and when Full transcript with speaker labels VibeVoice-ASR, AssemblyAI, Deepgram, Speechmatics
Speaker ID / Verification Is this speaker Alice? Identity label or similarity score CAM++, ECAPA-TDNN, pyannote voiceprint

This report focuses on the first two. The third is usually a pre- or post-processing component — segment first, then use voiceprints to decide who it was, or to replace “SPEAKER_00” with a real name.

2. The Global Landscape

2.1 Open Source

Option Org Key DER Speaker cap Streaming License
pyannote.audio 4.0 · Community-1 pyannoteAI (France) AISHELL-4 11.7% / AMI-IHM 17.0% Unlimited Via commercial Live-1 Code MIT / weights CC-BY-4.0
pyannote.audio 3.1 (legacy) pyannoteAI (France) AISHELL-4 12.2% / AMI-IHM 18.8% Unlimited ✗ MIT
NeMo Sortformer v2 NVIDIA DIHARD3 14.63% (≤4 spk) / CALLHOME-2 6.27% 4 (hard cap) ✓ (214× real-time) CC-BY-NC-4.0
NeMo MSDD (cascaded) NVIDIA DIHARD3 29.40% / CALLHOME-2 11.41% Can exceed 4 ✓ CC-BY-NC-4.0
3D-Speaker (CAM++ / ERes2NetV2) Alibaba Tongyi (China) AISHELL-4 10.30% / AliMeeting 19.73% No hard cap Via FunASR Apache-2.0
FunASR (CAM++ + FSMN-VAD + spectral clustering) Alibaba (China) AISHELL-4 13.3% No hard cap ✓ WebSocket Apache-2.0
VibeVoice-ASR (7B / 8.3B) Microsoft MLC DER 4.28% / cpWER 11.48% Not published Separate streaming release Open weights
WhisperX + pyannote Community Inherits pyannote, plus alignment noise Unlimited ✗ BSD-2
SpeechBrain (ECAPA-TDNN) SpeechBrain Only comparable with oracle VAD Unlimited ✗ Apache-2.0
Kaldi (x-vector) Community Classic baseline — ✗ Apache-2.0

Lower DER is better. Protocol warning: these figures come from each model card’s strictest setting (no collar, overlap scored), but baselines and training data differ — compare magnitudes, not decimals.

2.2 Commercial Services

Service Price Accuracy claim Cap Deployment
pyannoteAI Precision-2 €0.112/hr (Dev) / €0.096 (Starter); €19/mo incl. 170 hr Lowest DER in all ten benchmark domains Unlimited Cloud + on-prem
pyannoteAI Live-1 Priced separately Streaming scenarios Unlimited Cloud
AssemblyAI Universal-3.5 Pro $0.21/hr + diarization $0.02/hr cpWER 30.17 (own benchmark) 20 async / 10 streaming Cloud
Deepgram Nova-3 $0.0043/min batch (≈$0.258/hr) + diarization $0.0020/min cpWER 37.92 EN (third-party figure) No hard cap Cloud + self-host
Speechmatics Ursa 2 ≈$1.20/hr Strong on hard audio, multilingual 2–20 configurable Cloud + on-prem
ElevenLabs Scribe v2 $0.22/hr batch / $0.39/hr realtime cpWER 35.26 (third-party figure) 32 Cloud
Gladia $0.61/hr (down to $0.20/hr on Growth) Backed by pyannoteAI Precision-2 — Cloud (EU/US)
Rev AI $0.02/min (human review $1.50/min) DER 10–13% AMI — Cloud
OpenAI gpt-4o-transcribe-diarize $0.36/hr No published DER Not published Cloud
Google Cloud Chirp 3 $0.016/min (dynamic batch $0.004/min) Trails on hard audio 8 Cloud (V2 API)
Azure AI Speech Realtime $0.0167/min, batch $0.006/min Trails Speechmatics 4–10 Cloud + container
AWS Transcribe $0.024/min Diarization scored 8.3/10 10 Cloud
Tencent Cloud (China) Role separation ¥0.85/hr; 1:N voiceprint ¥4.2/1k calls No published DER — Cloud
iFlytek Tingjian (China) Free 2 hr/mo; ¥29/hr or ¥199/mo 98%+ Chinese recognition — Cloud + offline

List prices as of September 2026, all billed by audio duration. For most vendors diarization is a paid add-on, not a base feature (Deepgram +$0.0020/min, AssemblyAI +$0.02/hr).

3. Deep Dive, One by One

pyannote.audio 4.0 · Community-1 Best open source Score 8.9

Source: pyannoteAI (France) · GitHub
License: Code MIT / weights CC-BY-4.0 (attribution required)
Architecture: segmentation + overlap-aware + speaker embeddings + VBx clustering
Requirements: Python 3.10+, PyTorch 2+, ffmpeg, HF access token

The de facto standard in open source. Community-1 replaces the older segmentation and embedding models and adds VBx clustering, which substantially improves speaker assignment and speaker counting — exactly the two metrics that hurt 3.1. It also adds an “exclusive” mode that emits only the single most likely speaker at any moment, purpose-built for aligning with Whisper-style word timestamps and eliminating the most annoying failure mode in cascaded pipelines. The library sees roughly 45 million downloads a month on Hugging Face; most commercial transcription products quietly run it underneath.

Strengths

  • Highest accuracy in open source: AISHELL-4 11.7% DER
  • Clustering architecture, no hard speaker ceiling
  • Language-agnostic — no special handling for Chinese or others
  • Exclusive mode solves ASR timestamp reconciliation
  • Self-host, hosted, or upgrade to Precision-2 with one parameter

Weaknesses

  • Needs a GPU and MLOps — self-hosting cost is operational
  • Hugging Face weights are gated; accept terms per model
  • Open pipeline covers batch only; streaming means commercial Live-1
  • Speaker counting still errs above 8 speakers

pyannoteAI · Precision-2 / Live-1 Accuracy ceiling Score 9.1

Source: pyannoteAI (France, founded 2020) · pyannote.ai
Price: €0.112/hr (Dev) / €0.096 (Starter)
Extras: voiceprints (€0.015 each), confidence scores, STT orchestration
Deployment: Cloud / on-prem (Enterprise)

On every publicly checkable benchmark, Precision-2 takes the lowest DER — including the vendor’s own ten-domain DIHARD benchmark (259 recordings, ~67 hours) and an independent academic study over 196.6 hours of multilingual audio (English, Mandarin, German, Japanese, Spanish) that measured 11.2% DER, best in field. Official figures put it ~28% more accurate than the open Community-1, and 2.2–2.6× faster when self-hosted. Its distinguishing feature is voiceprints: enrol once, then turn “SPEAKER_00” into a real name and recognise the same person across files. For multi-speaker video dubbing — where each character’s voice must map to a fixed speaker — that capability is worth a lot.

Strengths

  • Lowest DER in all ten domains — no weak spot
  • No speaker cap, overlap detection included
  • Voiceprints enable cross-file identity recognition
  • STT orchestration returns an attributed transcript in one call
  • On-prem option; customer audio never used for training

Weaknesses

  • Pricier than AssemblyAI / Deepgram by one tier
  • Benchmark is vendor-maintained; independent replication still limited
  • Streaming requires Live-1; no open-source equivalent
  • You still need to pair it with an ASR (or use its orchestration)

NVIDIA NeMo · Sortformer v2 King of a narrow envelope Score 7.1

Source: NVIDIA NeMo / Riva · Hugging Face
License: CC-BY-NC-4.0 (non-commercial)
Architecture: end-to-end Transformer with Sort Loss (arrival-order labels)
Training data: 2,445 hrs real conversation + 5,150 hrs simulated multi-speaker

Sortformer sidesteps the permutation problem by sorting speakers in arrival order, emitting labels end-to-end without clustering. Inside its envelope the numbers are hard: DIHARD-III (≤4 speakers) 14.63% versus 29.40% for the cascaded MSDD, CALLHOME-2 6.27% versus 11.41% — and it is fast, with the streaming version hitting 214× real-time. There is exactly one problem, and it is fatal: the model detects a maximum of four speakers. That is architectural, not a tuning default — the output layer is built for four. Performance degrades beyond that, and a 14-speaker recording is simply out of scope.

Strengths

  • Doubles MSDD accuracy inside the ≤4-speaker envelope
  • End-to-end single model — no clustering pipeline to assemble
  • Native overlapped-speech detection
  • Streaming version at 214× real-time, very low latency
  • Tight Riva SDK integration for enterprise stacks

Weaknesses

  • Four speakers maximum — breaks on meetings, panels, call transfers
  • Trained primarily on English; mediocre on Chinese
  • CC-BY-NC-4.0 forbids commercial use
  • NeMo is a research framework; Hydra configs are a learning curve

Alibaba 3D-Speaker + FunASR Best for Chinese Score 8.4

Source: Alibaba Tongyi / DAMO Academy (China) · 3D-Speaker
License: Apache-2.0 (commercial use, no extra terms)
Core models: CAM++ (7.2M params, 0.65% EER on VoxCeleb1-O), ERes2NetV2
Pipeline: FSMN-VAD + embeddings + spectral clustering, WebSocket streaming

The strongest open-source answer for Chinese audio. 3D-Speaker reaches 10.30% DER on AISHELL-4 (Chinese meetings) — better than pyannote 3.1’s 12.2% and Community-1’s 11.7%; CAM++ hits 0.65% EER on VoxCeleb1-O with only 7.2M parameters, which is remarkable value. It is also the only public toolkit supporting true multimodality (audio + visual + semantic), letting mouth shape and on-screen presence assist who-is-speaking decisions. FunASR packages CAM++, FSMN-VAD and spectral clustering into one pipeline — the highest level of integration for Chinese production deployments — with a Docker image that serves streaming speaker labels over WebSocket.

Strengths

  • Leads on Chinese audio: AISHELL-4 10.30% DER
  • Apache-2.0 — cleanest commercial terms of any option here
  • Spectral clustering, no hard speaker cap
  • Streaming speaker labels supported
  • Multimodal extension (lip / visual cues)

Weaknesses

  • Dependency conflicts are common (fastcluster / hdbscan / libnvrtc)
  • Weaker than pyannote on English and multilingual audio
  • Documentation is Chinese-first; thin overseas community support
  • Unstable on English benchmarks like VoxConverse

Microsoft VibeVoice-ASR / Streaming The unified paradigm Score 8.3

Source: Microsoft Research (Furu Wei’s team) · Hugging Face
Size: ~8.3B (≈17.3 GB in BF16)
Capability: ASR + speakers + timestamps in one pass; up to 60 min per request
Languages: 50+, native code-switching

The most important architectural signal of 2026. The offline version ingests up to 60 minutes of audio without chunking and emits structured JSON (speaker ID + start/end + text), which removes the “speaker drift” that chunked pipelines introduce at every boundary; on the MLC-Challenge multi-speaker benchmark it posts DER 4.28% and cpWER 11.48%. The streaming version goes further — among the first LLM-based end-to-end streaming speaker-attributed ASR systems, emitting “who said what” as audio arrives, with an expected speaker-attribution latency of 2.00 seconds versus a measured 8.21 s for Azure ConversationTranscriber and 9.12 s for Google Cloud STT. It takes best or tied-best in 12 of 13 speaker-attribution settings.

Strengths

  • Unified model removes cascaded alignment error
  • 2.0 s attribution latency — 4× faster than cloud vendors
  • 60-minute single pass, best long-form consistency
  • 50+ languages, native Chinese-English switching
  • Custom hot-words; deployable via vLLM

Weaknesses

  • 7B/8.3B needs 17 GB+ VRAM — high barrier
  • General ASR WER 7.77% — good, not class-leading
  • Speaker-count ceiling not publicly documented
  • Streaming loses 5–6.7 cpWER points versus offline

WhisperX + pyannote Most practical combo Score 8.3

Source: Community (Max Bain et al.) · GitHub
License: BSD-2
Composition: Whisper transcription + wav2vec forced alignment + pyannote diarization
Output: Word-level timestamps + speaker labels

The engineering wrapper that stitches Whisper, forced alignment and pyannote into a single command — the default answer to “I want it running today”. Its value is word-level speaker labels: not just “this segment is speaker 00”, but every word mapped back to a speaker, which subtitles and line-by-line dubbing require. The cost is two layers of error stacking: alignment adds noise, and on real messy audio (Scribie’s production set) it measures 26.31% DER — not meaningfully better than AssemblyAI’s 26.61%.

Strengths

  • One command to run; the deepest pool of tutorials and answers
  • Word-level timestamps + speaker labels, ideal for subtitles
  • Whisper covers 99 languages
  • Completely free, offline, air-gap capable

Weaknesses

  • ASR and diarization errors compound
  • Real messy audio can reach 26%+ DER
  • Batch only, no streaming
  • Whisper itself struggles on overlapped multi-speaker audio

AssemblyAI · Universal-3.5 Pro Least integration effort Score 8.6

Price: $0.21/hr (async) + diarization $0.02/hr · assemblyai.com
Cap: 20 speakers async, 10 streaming
Metric: cpWER 30.17 (vendor comparison)
Extras: Speaker Identification maps labels to roles

The best developer experience in the category: one key, no tuning, and you get a speaker-labelled transcript plus summarisation and entity detection. Universal-3.5 Pro, released in July 2026, optimises for cpWER (which ties speaker labels to the actual transcribed words) and scores 30.17 in the vendor’s own comparison, ahead of ElevenLabs Scribe v2 (35.26), Gladia (36.87) and Deepgram Nova-3 EN (37.92). Speaker Identification can replace “Speaker A” with role names like “Agent” or “Customer” — but note that every one of those numbers comes from the vendor’s own test set.

Strengths

  • Lowest integration cost — one API key
  • 20 speakers async / 10 streaming is plenty for most products
  • Role-name labels instead of generic speaker IDs
  • Realtime and batch on the same interface
  • Native code-switching across 18 languages

Weaknesses

  • cpWER 30.17 is a vendor figure with no third-party replication
  • Cloud only — no self-hosted or on-prem option
  • Diarization is an ASR feature; nothing to tune when labels are wrong
  • Billed by session duration, so short audio costs disproportionately more

Deepgram · Nova-3 / Flux Speed and cost leader Score 8.7

Price: $0.0043/min batch (≈$0.258/hr) · deepgram.com
Cap: No hard speaker limit
Latency: <300 ms streaming; Flux has built-in turn detection
Deployment: Cloud + self-hosted

If what you want is “cheap and fast”, this is essentially the only answer. Nova-3 emits words and speaker labels from a single forward pass, is language-agnostic, has no hard speaker cap, and has the lowest batch unit price among major APIs. Flux, shipped in late 2025, goes further by folding turn detection (whose turn is it next) into the recogniser itself, removing the need for an external VAD — a structural advantage for voice agents. The trade-off is that diarization is a feature of the ASR engine, so on hard audio (crosstalk, overlap, heavy accents) accuracy visibly falls behind.

Strengths

  • Lowest unit price: $0.0043/min batch
  • Sub-300 ms latency, best streaming experience
  • No hard speaker ceiling
  • Flux builds in turn detection — voice-agent friendly
  • Self-hosting available

Weaknesses

  • cpWER 37.92 — last among the three-way comparison
  • Accuracy drops noticeably on hard multi-speaker audio
  • 36+ languages, fewer than the hyperscalers
  • Diarization and several enrichments are billed separately

Speechmatics · Ursa 2 Compliance and sovereignty Score 8.2

Price: ≈$1.20/hr (enterprise tier) · speechmatics.com
Cap: 2–20 configurable
Deployment: Cloud + on-prem / air-gapped
Languages: 50+, strong code-switching

The answer for hard audio and compliance requirements. It has the strongest reputation on accented English and on domains where a mislabelled speaker costs money — legal, medical, broadcast — and is one of the few commercial services offering on-prem and air-gapped deployment, meaning audio never leaves the internal network. In several regulated industries that single line decides procurement. The price is among the highest of the mainstream APIs (≈$1.20/hr, about 4.6× Deepgram) and it sells through enterprise channels, which rules out casual pilots.

Strengths

  • Best reputation on hard audio (accents, crosstalk, multilingual)
  • On-prem and air-gapped deployment
  • Configurable 2–20 speakers
  • Enterprise SLAs and compliance tooling

Weaknesses

  • Most expensive mainstream option — ~4.6× Deepgram
  • Enterprise sales only; high barrier to trial
  • All accuracy rankings are vendor self-reported
  • Smaller ecosystem than the hyperscalers

The Cloud Trio · Azure / AWS / Google Already on your bill

Azure: Azure AI Speech — realtime $0.0167/min, batch $0.006/min; Fast Transcription exposes maxSpeakers
AWS: Amazon Transcribe — $0.024/min; up to 10 speakers; channel identification is cleaner
Google: Cloud Speech-to-Text — $0.016/min (dynamic batch $0.004/min); up to 8 speakers; Chirp 3 is V2-API only
Shared weakness: diarization accuracy trails dedicated vendors

These three are valuable not for being best but for already being on your invoice. If your audio sits in S3, Blob Storage or GCS, they save you a data hop. But all three treat diarization as an accessory of a general speech platform: Google Chirp 3 caps at 8 speakers, AWS at 10, and Azure’s Fast Transcription recommends 4 (its multi-device conversation mode reaches roughly 10). There is also a commonly missed shortcut — if the recording is dual-channel (agent and customer on separate tracks), use channel identification rather than diarization: the channel already tells you who is who, no inference required, and it is free of the usual error modes.

Strengths

  • Zero friction with existing cloud infrastructure
  • Widest language coverage (Azure 140+, Google 125+)
  • AWS and Azure offer HIPAA BAA paths
  • Batch pricing as low as $0.004/min

Weaknesses

  • Low speaker caps (Google 8, AWS 10)
  • Accuracy trails dedicated vendors on hard multi-speaker audio
  • Almost nothing tunable in the diarization path
  • Vendor lock-in raises migration cost

4. Head-to-Head Matrix

Option Architecture Speaker cap Streaming Languages On-prem Commercial license
pyannote Community-1 Cascaded + VBx Unlimited ✗ Agnostic ✓ CC-BY-4.0
pyannoteAI Precision-2 Cascaded + voiceprints Unlimited ✓ (Live-1) Agnostic ✓ (Enterprise) Commercial
NeMo Sortformer v2 End-to-end 4 ✓ English-first ✓ Non-commercial
3D-Speaker / FunASR Cascaded + spectral No hard cap ✓ (FunASR) Chinese-first ✓ Apache-2.0
VibeVoice-ASR End-to-end LLM Not published ✓ (streaming release) 50+ ✓ Open weights
WhisperX + pyannote Cascaded + alignment Unlimited ✗ 99 ✓ BSD-2
AssemblyAI U-3.5 Pro Unified ASR 20 / 10 ✓ 18 (99 via U-2) ✗ Commercial
Deepgram Nova-3 Unified ASR No hard cap ✓ 36+ ✓ Commercial
Speechmatics Ursa 2 Unified ASR 2–20 configurable ✓ 50+ ✓ Commercial
Azure / AWS / Google Unified ASR 4–10 ✓ 100+ Azure container only Commercial

5. Measured Data: The Part Most Write-ups Soften

This section does one thing: puts numbers back inside the protocol they came from. Diarization is one of the few fields where a single model can be reported with a dozen different DER values, so a comparison without a protocol note means nothing.

5.1 The Power of Protocol: One Model, Three Sets of Numbers

Model AISHELL-4 AMI-IHM CALLHOME VoxConverse DIHARD-III
pyannote 3.1 (legacy) 12.2% 18.8% 28.5% 11.2% 21.4%
pyannote Community-1 11.7% 17.0% 26.7% 11.2% 20.2%
pyannoteAI Precision-2 11.4% 12.9% 16.6% 8.5% 14.7%
3D-Speaker 10.30% 21.76% (SDM) — 11.75% —
NeMo Sortformer v2 ~13% ~13% 6.27% (2 spk) — 14.63% (≤4 spk)

Protocol: pyannote figures use the strictest setting — no collar, overlap scored. Two things stand out. (1) Sortformer’s 6.27% on two-speaker CALLHOME is the best number in the table, but the same model caps at four speakers on DIHARD-III — its strength is “few speakers, clean”. (2) 3D-Speaker beats every pyannote version on Chinese meetings (10.30%) yet falls to 21.76% on English AMI-SDM. Language specialisation cuts both ways.

5.2 Your Room Sets the Ceiling: Microphones Beat Models

Audio type Benchmark pyannote 3.1 DER Note
Broadcast (close-mic) REPERE 7.8% Studio-grade audio; every system looks good
Web video VoxConverse 11.3% Political debate; overlap present, quality OK
Chinese meetings AISHELL-4 12.2% Multi-party, far-field capture
Close-mic meetings AMI (IHM) 18.8% One mic per person, but frequent natural turns
Mixed hard domains DIHARD-III 21.7% Clinical, courtroom, children’s speech
Far-field meetings AliMeeting 24.4% Single room mic, heavy reverb and crosstalk
In-the-wild video AVA-AVD 50.0% Music, crowds, heavy overlap — half the time mislabelled

From pyannote’s official model card (strictest setting). The conclusion is harder than any buying advice: your diarization quality is set first by microphones and acoustics, and only then by the model. The same model scores 7.8% on broadcast audio and 50% on in-the-wild video. If the budget is finite, buy one close mic per speaker before you buy a better model.

5.3 The Cost and Benefit of Unification: Attribution Latency

System Speaker-attribution latency Five-set average error Type
VibeVoice-ASR-Streaming 2.00 s 24.66 End-to-end LLM, streaming
Gemini 3.5 Transcribe Live — 25.23 Cloud streaming
Azure ConversationTranscriber 8.21 s — Cloud streaming
Google Cloud STT 9.12 s — Cloud streaming
ElevenLabs Scribe v2 Realtime — 41.39 Cloud streaming

Conditions: Microsoft VibeVoice-ASR-Streaming technical report (arXiv:2609.02812), 22-frame chunks with 4-frame lookahead; evaluation audio truncated to 8 minutes across 10 languages. The key difference is not accuracy but how soon the system dares to commit: cloud vendors revise speaker labels tens of seconds later, while the end-to-end model settles within 2 seconds. For live captioning and voice agents, that gap matters more than a point of DER.

5.4 Real Messy Audio: How Far Lab Numbers Fall

System N-WER (word errors) DER (speaker errors)
WhisperX (large-v3 + pyannote) 12.81% 26.31%
AssemblyAI 15.13% 26.61%
Deepgram 15.62% 27.23%

Conditions: Scribie production set — 26 real transcription jobs including legal depositions, multi-speaker interviews, compressed Zoom audio, phone recordings and terminology-dense content. This is the pair of numbers to remember: models that score 11% DER on podcast-clean audio commonly land at 26–27% on real call and deposition audio. Note also that DER is universally worse than word error rate — attributing speech to a speaker is harder than hearing the words — and that all three systems effectively tie here, which says the ceiling belongs to the task, not to the product.

6. Scoring and Head-to-Head Verdicts

The previous sections analysed each option qualitatively; this one puts numbers on it. This scoring is not an official benchmark — it is a composite judgment drawn from public model cards, vendor documentation, independent academic evaluation and production measurements. Seven dimensions are each scored out of 10 and weighted into a single overall score. Results will vary with audio conditions and deployment model, so treat this as a starting point for ranking, not the final word.

6.1 Scoring Dimensions and Weights

Dimension Weight What it measures
Accuracy 30% DER / cpWER on public benchmarks, plus stability on real messy audio
Speaker capacity 10% Whether the speaker ceiling is hard or tunable; failure at high speaker counts
Language coverage 15% Language-agnostic design, Chinese performance, code-switching
Real-time capability 10% Streaming support, attribution latency, whether an external VAD is needed
Usability 10% Integration cost, dependency complexity, docs, whether you must wire up your own ASR
Cost 15% Effective per-hour price, or the GPU and engineering cost of self-hosting
Licensing & compliance 10% Clarity of commercial terms, on-prem availability, training-data commitments

6.2 Overall Scoreboard

Rank Option Overall score Stars One-line verdict
1 pyannoteAI Precision-2
commercial
9.1
★★★★★ Wins all ten benchmark domains; no capacity or compliance gap — you just pay
2 pyannote Community-1
open source
8.9
★★★★★ Best open source, upgradeable to the commercial model — you pay in ops
3 Deepgram Nova-3
commercial
8.7
★★★★★ First on speed and first on price; accuracy is what it trades away
4 AssemblyAI U-3.5 Pro
commercial
8.6
★★★★★ Easiest to integrate, best unified transcript UX; accuracy awaits third-party proof
5 Alibaba 3D-Speaker / FunASR
open source · China
8.4
★★★★★ Best for Chinese and fully Apache-2.0; environment setup is the barrier
6 WhisperX + pyannote
open-source pipeline
8.3
★★★★★ Practical one-command combo; you pay in compounded error
7 VibeVoice-ASR
open source · Microsoft
8.3
★★★★★ Most advanced architecture, leading long-form and streaming attribution; heavy VRAM cost
8 Speechmatics Ursa 2
commercial
8.2
★★★★★ The compliance and hard-audio answer; price is the barrier
9 NeMo Sortformer v2
open source · NVIDIA
7.1
★★★★★ Superb at ≤4-speaker English, but the 4-speaker cap and non-commercial licence hold it down

Overall score = sum of dimension scores × weights (out of 10). Ranks 6 and 7 tie at 8.3 — WhisperX wins on usability and ecosystem, VibeVoice on architecture and long-form consistency.

6.3 Seven-Dimension Breakdown

Option Accuracy Capacity Language Real-time Usability Cost Compliance Overall
Precision-2 9.8 10.0 9.5 9.0 8.5 7.0 9.0 9.1
Community-1 9.3 10.0 9.5 6.5 6.5 9.5 9.5 8.9
Deepgram Nova-3 8.0 10.0 8.0 10.0 9.5 9.5 7.5 8.7
AssemblyAI U-3.5 8.5 8.5 8.5 8.5 10.0 8.5 8.0 8.6
3D-Speaker / FunASR 8.5 9.5 7.5 7.5 6.0 9.5 10.0 8.4
WhisperX + pyannote 7.8 10.0 9.5 5.0 7.0 9.5 9.0 8.3
VibeVoice-ASR 8.5 8.0 8.5 9.0 7.5 8.0 8.5 8.3
Speechmatics Ursa 2 8.5 8.0 9.0 8.0 8.0 6.0 9.5 8.2
NeMo Sortformer v2 8.5 3.0 6.0 9.5 6.0 9.0 4.0 7.1

Green marks the top tier in each column, amber a clear weakness, red the weakest. Three signals. (1) Sortformer’s accuracy (8.5) and capacity (3.0) differ by 5.5 points — it is not a weak model, it is a strong model locked inside an envelope. (2) Every open-source option loses in the “real-time” column: streaming has long been a commercial-side capability only. (3) Compliance separates Sortformer on its own — a non-commercial licence is a veto in commercial settings.

6.4 Category Champions

Accuracy champion
pyannoteAI Precision-2

Lowest DER in all ten DIHARD domains; 11.2% best-in-field in an independent academic study across 196.6 hours and five languages.

Open-source accuracy
pyannote Community-1

AISHELL-4 11.7%, AMI-IHM 17.0% — ahead of most commercial services on public benchmarks.

Chinese audio
Alibaba 3D-Speaker

10.30% DER on AISHELL-4, better than any pyannote release; CAM++ reaches 0.65% EER with 7.2M parameters.

Speaker capacity
pyannote / Deepgram

Clustering architectures have no hard speaker ceiling; Sortformer’s 4, Google’s 8 and AWS’s 10 are architectural limits.

Real-time
Deepgram / VibeVoice-Streaming

Deepgram streams under 300 ms; VibeVoice attributes speakers in 2.0 s — four times faster than Azure.

Cost
Deepgram / self-hosted Community-1

$0.0043/min is the commercial floor; self-hosting Community-1 costs one modern NVIDIA GPU.

Licensing
3D-Speaker / FunASR

Apache-2.0 with no extra terms; community options sit between “attribute us” and “no commercial use”.

Sovereign deployment
Speechmatics / pyannoteAI Enterprise

On-prem and air-gapped options keep audio inside the network; Azure containers are the runner-up.

6.5 Head-to-Head Verdicts: Four Matchups

① Accuracy: three protocols, three winners

pyannoteAI claims all ten domains on DER; AssemblyAI claims 30.17 cpWER against Deepgram’s 37.92; Speechmatics claims 25% ahead of its nearest competitor. All three are true and all three are partial — each picked the metric and test set that favours it. The only robust conclusion: dedicated diarization models (the pyannote family) reliably beat diarization-as-an-ASR-feature on hard audio.

② Open source: three constraint lines

Community-1 wins on accuracy + no speaker ceiling; 3D-Speaker wins on Chinese + Apache-2.0; Sortformer wins on speed (214× real-time) but is locked by the 4-speaker cap and a non-commercial licence. Ask three questions first — how many speakers, which language, commercial or not — and two of the three are eliminated.

③ Architecture: unified vs cascaded

Cascaded (WhisperX + pyannote) wins on flexibility, swappable parts and word-level timestamps, but compounds two layers of error: 26.31% DER on real messy audio. Unified (VibeVoice / AssemblyAI) wins on eliminating alignment error and attributing in 2.0 s versus 8–9 s, at the cost of a black box you cannot tune when it errs.

④ Cost: the self-host breakeven

By published figures, self-hosting only beats API pricing above roughly 14,000 audio hours per month (including GPU and engineering). Below that, hosted APIs are cheaper — but if audio cannot leave your network, self-hosting stops being a choice and becomes the only option. pyannoteAI’s €0.035/hr hosted open-source tier is the middle path: no GPU, no model lock-in.

Verdict summary: there is no all-round champion, only a best fit per scenario —

  • Accuracy first, budget available, on-prem needed: pyannoteAI Precision-2 (9.1)
  • Open source preferred, ops accepted: pyannote Community-1 (8.9)
  • Speed and lowest unit price: Deepgram Nova-3 (8.7)
  • Zero ops, fastest to ship: AssemblyAI Universal-3.5 Pro (8.6)
  • Chinese audio with Apache-2.0 clarity: Alibaba 3D-Speaker / FunASR (8.4)
  • Word-level labels for subtitles and dubbing: WhisperX + pyannote (8.3)
  • Long-form single-pass and streaming attribution: VibeVoice-ASR (8.3)
  • Hard compliance and sovereignty: Speechmatics (8.2)
  • English ≤4 speakers, efficiency only (non-commercial): NeMo Sortformer v2 (7.1)

7. Engineering Practice: Six Steps That Decide the Outcome

The model is one link in the chain. Miss any of the following and an 11% DER becomes 30%.

Step What it decides Practical guidance
Microphones and capture The accuracy ceiling One close mic per speaker beats a single room mic. The same model scores 7.8% close-mic and 50% in the wild — no model recovers lost room acoustics
VAD preprocessing False alarms Run FSMN-VAD or MarbleNet first to strip music, noise and silence; false alarms fall markedly
Overlap handling Success on crosstalk Ranking: Sortformer (native) > pyannote (overlap-aware) > pure clustering. For meetings, favour models that handle overlap
Speaker-count priors Counting errors Always pass num_speakers when known; when unknown, give min/max bounds. Both beat letting the model guess
Voiceprint enrolment IDs to real names Build a voiceprint store with CAM++ or ECAPA-TDNN to map SPEAKER_00 to a real identity; pyannoteAI voiceprints are the managed route
ASR reconciliation Word-level label quality Use pyannote’s exclusive mode so only one speaker is active per moment, preventing overlap regions from misattributing words

The other half of the cost story: the self-hosting bill is not just GPU. One modern NVIDIA card runs Community-1, but you also own model version updates, expired Hugging Face tokens (which break the whole pipeline silently), dependency conflicts (3D-Speaker has long been stuck on fastcluster / libnvrtc), and batch scheduling. Hosted APIs are selling those, not just the model.

A shortcut people miss: if the recording is dual-channel (agent and customer on separate tracks), use channel identification instead of diarization. The channel tells you who is who with zero inference and zero error rate. AWS and Azure both build this in; many teams pay for diarization they never needed.

8. Recommendations by Scenario

Meeting transcription, unknown headcount

pyannote Community-1

No speaker ceiling, language-agnostic — holds from 5 to 15 participants. One GPU, or hosted at €0.035/hr.

Call centre QA and conversation analytics

Deepgram Nova-3 or pyannoteAI Precision-2

Dual-channel calls need channel identification, not diarization. For single-channel crosstalk precision, go Precision-2.

Chinese meetings and podcasts

Alibaba 3D-Speaker + FunASR

10.30% DER on AISHELL-4 is the best in the field, Apache-2.0 for commercial use, streaming labels included.

Live captioning and voice agents

Deepgram Flux or VibeVoice-Streaming

The former builds in turn detection under 300 ms; the latter attributes speakers in 2.0 s, four times faster than cloud vendors.

Video dubbing and multi-character translation

pyannoteAI Precision-2 + voiceprints

Voiceprints bind characters to fixed voices, then the diarization output feeds TTS for multi-character dubbing.

Subtitles needing word-level labels

WhisperX + pyannote

The only out-of-the-box combination that yields word timestamps plus speakers — SRT straight out.

Long-form single-pass (60 minutes)

VibeVoice-ASR (offline)

No chunking, no stitching, no speaker drift; structured JSON with timestamps out.

Audio cannot leave the building

Speechmatics on-prem / pyannoteAI Enterprise

Both support air-gapped deployment; in China, self-hosted 3D-Speaker is the equivalent route.

Tightest budget, just get started

FunASR or WhisperX

Fully open source with zero API cost on a consumer GPU; validate on a small sample before scaling.

Decision Framework

1. Count speakers first: ≤4 and English → Sortformer works (mind the licence); 5–20 → 3D-Speaker or pyannote; 20+ → clustering only (pyannote, Deepgram)

2. Then language: Chinese-first → 3D-Speaker / FunASR; multilingual → the pyannote family (agnostic); English-only → widest choice

3. Can audio leave your network? No → self-host or Speechmatics / pyannoteAI Enterprise; Yes → hosted APIs save the ops

4. Commercial use? Yes → rule out Sortformer (non-commercial), honour CC-BY-4.0 attribution, or pick Apache-2.0 for the cleanest terms

5. Need real-time? Yes → Deepgram (fastest) or VibeVoice-Streaming (fastest attribution); No → batch saves 3–6×

6. Is the recording dual-channel? Yes → use channel identification first and skip the diarization bill

7. Need word-level labels? Yes → WhisperX + pyannote, or pyannote’s exclusive mode

8. Still undecided? → go by the overall scores: Precision-2 (9.1) for accuracy, Community-1 (8.9) for open source, Deepgram (8.7) for speed and cost

9. Pitfalls and Compliance

1. Always quote DER with its protocol. An 11% with no collar and overlap scored is not the same number as an 11% with a collar and overlap ignored. Look for those two switches before reading any comparison table — vendors routinely report the flattering side.

2. Establish the speaker ceiling early. Sortformer’s 4, Google’s 8, AWS’s 10 and Azure’s 4–10 are architectural limits, not tuning issues. One extra participant invalidates the whole approach — eliminate this risk before the POC, not after.

3. A non-commercial licence is a veto. NeMo Sortformer weights are CC-BY-NC-4.0 and cannot be used commercially. Teams that ship it into production and discover this at legal review face very expensive rework. For commercial projects choose Apache-2.0 (3D-Speaker / FunASR) or an explicit commercial licence.

4. CC-BY-4.0 requires attribution. pyannote Community-1 weights are commercially usable with an attribution notice in your product. Code under MIT and weights under CC-BY are two separate licences — do not read only the code half.

5. Read the training-data clause. pyannoteAI states customer audio is never used for training and publishes retention policy per tier; most vendors make no such commitment. For sensitive audio, put this in the contract.

6. Never trust a single vendor benchmark. Three vendors each claiming first place is normal market behaviour. The correct move: take 20–50 hours of your own real audio (including the worst of it), score every candidate with the same pyannote.metrics protocol, then decide.

7. Do not neglect VAD and silence. False alarms — noise counted as speech — are one of the three components of DER. A clean VAD in front often helps more than a model upgrade.

8. Voiceprints trigger biometric-privacy law. Using voiceprints to identify people brings biometric data rules into play: sensitive personal information with separate-notice consent obligations in China, special-category data under GDPR in the EU, and class-action exposure under BIPA-style US state laws. Clear it with legal before shipping identity features.

Key Findings

1. The field is moving from a separation module to a unified capability. The 2026 dividing line is speaker-attributed ASR — VibeVoice compresses ASR, diarization and timestamps into one forward pass; AssemblyAI and Deepgram build labels into the recogniser. The cascaded pipeline’s alignment error and latency disadvantage will only grow, though its swappable parts remain the right answer for some products.

2. The accuracy ceiling belongs to the pyannote family, and open source sits one parameter away from commercial. Community-1 leads open source, Precision-2 leads every benchmark domain, and both share one integration — a rare “validate free, upgrade seamlessly” path.

3. Chinese audio has its own winner. 3D-Speaker posts 10.30% DER on AISHELL-4, better than any pyannote release, under a fully commercial Apache-2.0 licence. Language specialisation cuts both ways — it drops to 21.76% on English AMI-SDM.

4. The speaker ceiling is the first hard filter. Four, eight and ten are not performance differences; they are the line between usable and unusable. Count your speakers before you read the accuracy table.

5. Microphones decide more than models do. The same model scores 7.8% on broadcast audio and 50% in the wild. When budget is finite, spend it on capture first.

6. Real audio is far harsher than any leaderboard. Production measurements cluster around 26–27% DER, and speaker errors are consistently worse than word errors. Any POC must run on your own messy data; clean-sample numbers carry no information.

7. Licensing is the buried landmine. Sortformer’s non-commercial terms, CC-BY attribution duties and biometric-privacy obligations around voiceprints can each stop a technically successful project before launch. Put them on the same table as accuracy when you decide.

Put Multi-Speaker Audio to Work

Diarization is the easy half; the hard half is keeping every voice intact across languages. DeepVideo by DeepForgeHub translates any video into 30+ languages with Voice Clone and lip-sync — a Windows & Mac desktop app that runs locally, so your footage never leaves your machine.

Try DeepVideo →

Free tier: 18 minutes total + 2 minutes daily · Local processing, no cloud upload

© 2026 DeepForgeHub Research. Sources: pyannoteAI public benchmarks and model cards, NVIDIA NeMo model cards and technical reports, Alibaba 3D-Speaker / FunASR documentation, Microsoft’s VibeVoice-ASR technical report (arXiv:2609.02812), official documentation and pricing pages from AssemblyAI, Deepgram, Speechmatics, ElevenLabs, Gladia and Rev AI, AWS / Azure / Google Cloud documentation, and third-party production benchmarks (Scribie, Neosophie). Prices are list prices; DER protocols are noted inline. As of September 2026.

Newsletter Updates

Enter your email address below and subscribe to our newsletter

Leave a Reply

Your email address will not be published. Required fields are marked *