Lip-Sync Models: A Global Comparison
Re-dubbing · Talking-Head Generation · Joint Generation — September 2026
Lip accuracy · Visual fidelity · Hardware barrier · Licensing · Cost
Executive Summary
In 2026, “lip-sync” is no longer one thing — it is four unrelated tracks: re-dubbing an existing video (lip-sync: LatentSync, sync.so), generating a full performance from a single image (talking head / avatar: OmniHuman, HunyuanVideo-Avatar), generating video and audio in one forward pass (joint generation: MOVA), and end-to-end translation dubbing pipelines (HeyGen). Pick the wrong track and even the best model is money wasted.
The global landscape compresses into two sentences. On the open-source side, the deciding factor is the method, not the parameter count — four technical routes (GAN: Wav2Lip; single-step latent inpainting: MuseTalk; latent diffusion: LatentSync; 3DMM: SadTalker) directly determine speed, resolution, and quality ceilings. On the commercial side, the ceiling is sync.so’s sync-3 — native 4K, automatic occlusion detection, extreme angles — but it costs $6.4–8.0 per output minute. China’s answer is entirely different: Volcano Engine’s OmniHuman 1.5 delivers 1080P avatars at ¥1/second, at the cost of a 60-second audio cap, concurrency of 1, and single-image input only.
Key findings:
- Licensing is the sharpest knife in this field. The most-cited model, Wav2Lip, explicitly bans commercial use (trained on LRS2); VideoReTalking is CC BY-NC-SA; SadTalker’s weights inherit OpenRAIL-M restrictions; Tencent’s Sonic README states outright that commercial use must go through Tencent Cloud’s API. The only clean, commercially usable options are MuseTalk (MIT), LatentSync, and EchoMimic (Apache-2.0).
- LSE-C / LSE-D should not drive procurement. The two most-cited lip-sync metrics have been shown by multiple papers to correlate poorly with human judgment, to be sensitive to cropping and brightness, and to be beatable by models that score far above ground truth — plus the circular reasoning of “train with SyncNet, score with SyncNet.”
- The silence test is the best one-shot veto. If the mouth keeps moving without speech, the model is copying lip trajectories from visual context rather than truly listening. HighSync scores 0.93; MuseTalk only 0.68 and LatentSync 0.81.
- The open-source vs. commercial price gap is over 20×. Self-hosted LatentSync costs electricity; sync-3 costs $6.4–8.0/min — but self-hosting means 18GB VRAM, environment setup, and maintenance. The math only tilts toward self-hosting above roughly 30 minutes of output per month.
- Visual quality and lip accuracy are two separate curves. MuseTalk scores highest on human-evaluated video quality (4.34) but lowest on lip accuracy (3.14) — adversarial refinement makes the frame pretty, yet single-step generation without true temporal audio modeling makes the mouth untrustworthy.
1. First, Sort the Task: Four Different “Lip-Sync” Jobs
These four tasks are constantly conflated, but their inputs, outputs, and available models are entirely different. Step one of any selection is knowing which box you are in.
| Task | Input → Output | What gets changed | Representative models |
|---|---|---|---|
| Lip re-sync Lip-sync / dubbing |
Existing video + new audio → same video | Redraws only the mouth region; performance, camera moves, background all preserved | LatentSync, MuseTalk, Wav2Lip, sync.so sync-3 |
| Talking-head generation Avatar |
One still image + audio → new video | Generates the whole performance (expressions, head motion, half-body gestures) | OmniHuman 1.5, HunyuanVideo-Avatar, Wan2.2-S2V, EchoMimic |
| Joint AV generation | Text / image → video + synchronized audio | Video and audio aligned in a single inference, no “render first, dub later” | MOVA (OpenMOSS) |
| End-to-end localization dubbing | Finished video → multilingual video | Translation + voice cloning + lip-sync + timeline stretching, one pipeline | HeyGen, Rask AI, Vozo, DeepBrain AI Studios |
A common misjudgment: using a talking-head generation model to fix lip-sync on real footage. Such models re-imagine the entire face’s motion amplitude — over-generation for already-shot material. You asked for “move only the mouth”; you get “re-act the whole thing.” Conversely, a re-sync model cannot handle a still image: sync.so’s lipsync-2 / 2-pro explicitly do not support static image input (only sync-3 does image-to-video).
2. Global Landscape at a Glance
Open source first, then commercial. Both tables work as candidate shortlists.
2.1 Open-Source Solutions
| Solution | Organization | Method | Resolution | VRAM | Speed | License | Overall |
|---|---|---|---|---|---|---|---|
| LatentSync 1.6 | ByteDance (China) | Latent diffusion + SyncNet supervision | 512×512 (1.5: 256) | ~18GB (1.5: ~8GB) | Not real-time; seconds to minutes per clip | Apache-2.0 | 8.4 |
| MuseTalk 1.5 | Tencent (China) | Single-step 256-region latent inpainting | 256×256 (face) | ~8GB | Real-time, 30fps+ on V100 | MIT | 8.3 |
| EchoMimic V2 / V3 | Ant Group (China) | Half-body animation + gesture generation | 512-class | V3 runs on ~12GB | Moderate | Apache-2.0 | 8.2 |
| HunyuanVideo-Avatar | Tencent Hunyuan (China) | MM-DiT + emotion module + face-aware audio adapter | Up to 720p | 24GB official, 10GB optimized | Slow (96GB recommended) | Tencent Hunyuan Community License | 7.9 |
| Wan2.2-S2V-14B | Alibaba Tongyi (China) | Video base model + AdaIN/CrossAttn audio control | 480P–720P | Official rec: single 80GB card | Slow; supports minute-long video | Apache-2.0 (family) | 7.9 |
| LivePortrait | Kuaishou (China) | Implicit keypoints + stitching retargeting | 256×256 up (upscaleable) | 6–8GB (min 4GB) | ~12ms/frame on 4090 | Code MIT / weights disputed | 7.8 |
| Wav2Lip | IIIT Hyderabad (India) | GAN + sync discriminator | 96×96 (face) | ~6GB | Fast | Research / non-commercial only | 6.3 |
| SadTalker (table only) | Xi’an Jiaotong Univ. et al. | 3DMM coefficients + face rendering | Medium | 6GB | Medium; slow on long video | Code Apache / weights inherit upstream | 6.5 |
| VideoReTalking (table only) | Academic team | Three-stage: expression neutralize → lips → enhance | Medium | Medium | Slow | CC BY-NC-SA 4.0 (non-commercial) | 6.4 |
| Sonic (table only) | Tencent | Audio-driven portrait animation | High | Medium | Medium | Non-commercial; commercial via Tencent Cloud | 6.6 |
| InfiniteTalk (table only) | MeiGen AI (China) | Image → unlimited-length talking video | Medium | Runs on 12GB-class | Medium | Apache-2.0 | 7.2 |
| MultiTalk (table only) | MeiGen AI (China) | Multi-speaker conversation lip-sync | ~450p native | ~8GB | Medium | Apache-2.0 | 7.1 |
| MOVA (table only) | OpenMOSS (China) | Joint video + audio in one inference | ≤720p, ≤8 s per clip | ~48GB (~12GB offloaded) | Heavy | Apache-2.0 | 7.0 |
| HighSync (table only) | Academic team | 512 latent diffusion + leakage-proof design | 512×512 | High | Diffusion-class | Per repository | 7.4 |
VRAM and speed figures come from official docs and model cards; real measurements vary widely with resolution, frame count, and optimizations. “Overall” scores are explained in Section 6 (not vendor benchmarks).
2.2 Commercial Services
| Service | Price | Lip / quality claim | Main limitations | Deployment | Overall |
|---|---|---|---|---|---|
| sync.so sync-3 | $0.107–0.133/s (~$6.4–8.0/min) | Native 4K, auto occlusion detection, can open fully closed mouths | Watermarked free tier, no real-time | Cloud API | 8.9 |
| Volcano Engine OmniHuman 1.5 | ¥1/s (China); BytePlus $0.12/s | 1080P, multi-character, lips + emotion + gestures | Audio <60 s (≤15 s recommended), concurrency 1 | Cloud API (via Jimeng) | 8.7 |
| HeyGen | From $29/mo; audio-only dubbing $0.10/min, lip-synced $0.24/min | 175+ language end-to-end translation dubbing | Lip-sync targets avatar/translation, not arbitrary footage repair | Cloud + API | 8.6 |
| D-ID (table only) | $5.90/mo (Lite) / $29.99 (Pro) | Photo → talking avatar | Quality degrades past ~60 s | Cloud + API + streaming | 7.6 |
| Synthesia (table only) | From $29/mo; ~$2.9/min effective | 140+ languages, strong enterprise governance | Output leans “broadcast anchor” | Cloud + API | 7.8 |
| Runway Act-One (table only) | From $15/mo | Performance-driven + expression transfer | Deep features gated to top tiers; hard lip cases lose to specialists | Cloud | 7.5 |
| Kling LipSync (table only) | ~$0.014/s on fal.ai | Lip-sync for generated characters, cheap | For generated footage, not live-action repair | Cloud API | 7.7 |
| Hedra (table only) | From $15/mo | Strong character expressiveness (face + head) | “A performance around a photo,” not footage repair | Cloud | 7.4 |
| Baidu XiLing (table only) | ¥7,999/avatar/yr (1500 min incl.); Basic ~¥2,800/mo | Claims 98.5% lip accuracy; strong gov/enterprise presence | Heavy cost for SMBs | Cloud + on-prem | 7.5 |
| Tencent Zhiying (table only) | ¥3,999/avatar/yr (500 min) + ¥3,999/voice/yr | Avatar cloning from 3-minute footage | Consumer / WeChat-Channels ecosystem focus | Cloud | 7.4 |
| SenseTime Ruying (table only) | ¥3,598/avatar/yr (500 min) + ¥598/voice/yr | Strong expression drive | Voice quota overruns easily at ¥2/min | Cloud + on-prem | 7.5 |
| Silicon Intelligence (table only) | S-tier avatar ¥3,980/yr + E-tier voice ¥680 | Avatar from one photo, natural lips | Tiered quotas; E-tier lower quality | Cloud + on-prem | 7.3 |
| iFlytek Zhizuo (table only) | Membership from ¥45/mo; ¥3–6.7/min by duration | Two decades of ASR/TTS heritage | Body motion and 3D rendering weaker | Cloud | 7.2 |
Prices are public list prices as of September 2026; commercial metering varies wildly (per second, per credit, per avatar-year) and has been normalized where possible.
3. Deep Dives
3.1 LatentSync 1.6 Open-source re-dub quality ceiling Apache-2.0 Overall 8.4
It made diffusion-based lip-sync actually work. The paper’s core contribution is not architecture but fixing the shortcut problem in diffusion lip models — the network cheats by copying mouths from neighboring frames instead of listening. The authors used SyncNet supervision to push discriminator accuracy from 91% to 94%, plus TREPA temporal alignment to suppress flicker. Version 1.6 merely retrained at 512×512; the architecture is unchanged, so one codebase serves both checkpoints — swap the checkpoint and the resolution parameter in the U-Net config.
In public academic comparisons it is the diffusion camp’s dual winner for quality and lip accuracy (LSE-C of 8.05 on HDTF, the top tier among compared methods) — but the price is blunt: not real-time, and 1.6 wants 18GB VRAM, which excludes most consumer GPUs. Use 1.5 to save money; go 1.6 when teeth and lip lines must be sharp.
Strengths
- Apache-2.0 — among the cleanest commercial licenses
- 512×512 output; teeth and lip detail clearly ahead of GAN-based rivals
- TREPA-treated temporal stability; far less flicker than the Wav2Lip era
- Specific optimizations for Chinese video since 1.5
- Well community-validated; hosted endpoints available off the shelf
Weaknesses
- 1.6 needs ~18GB VRAM; the 1.5→1.6 quality jump is a step function
- Not real-time; slow iteration; long videos take minutes
- Re-dubs existing video only; no static-image generation
- Main repo is less active — new scenarios mean DIY
3.2 MuseTalk 1.5 The only truly real-time one MIT Overall 8.3
Its trade-off is crystal clear: abandon iteration, buy speed. Because it never denoises in multiple steps, it is the only option on this list that can carry real-time conversation, and its video quality and identity preservation are arguably the best in class (FID 6.52, CSIM 0.86 on HDTF). But its lip-sync score (LSE-C) is visibly below the diffusion rivals — the canonical case of “quality 4.34 (highest), lips 3.14 (lowest)” in human evaluation.
The silence test matters even more: MuseTalk scores just 0.68 — the mouth keeps moving without speech, meaning it infers lip trajectories from visual context rather than truly relying on audio. For livestreaming (“the person is always talking”) this is survivable; for edited footage with pauses and silence, it shows.
Strengths
- MIT license — the least friction for commercial use
- Truly real-time (30fps+); the only open-source option that fits live pipelines
- Runs on 8GB VRAM; low hardware barrier
- Top-tier open-source video quality and identity preservation
- Stable across languages including Chinese
Weaknesses
- Face region capped at 256×256; 1080p requires external upscaling
- Lip accuracy clearly below diffusion rivals like LatentSync
- Silence test 0.68 — the mouth moves through silent passages
- Inter-frame jitter; long videos need extra temporal smoothing
3.3 Wav2Lip Most cited — and commercially forbidden Community model zoo Overall 6.3
It defined the field — and it is the most cited and most commercially abused model. Search “free lip-sync” today and most top results are third-party wrappers or web apps of Wav2Lip — while the original repo’s license plainly says non-commercial. Worse, dozens of forks quietly redistribute the same checkpoints, which does not constitute a license transfer.
Technically it has been surpassed across the board: 96×96 generation resolution means teeth and lip lines are inherently blurry, and no amount of “upscaled to 1080p” invents detail that was never generated. Its remaining value is validation and education: cheap, fast, low-VRAM — good for proving out a pipeline and learning to read LSE metrics. For actual commercial delivery, switch to MuseTalk / LatentSync / EchoMimic.
Strengths
- Lowest hardware bar — 6GB VRAM suffices
- Rich tutorials, wrappers, and ComfyUI nodes
- Native LSE metric implementation; handy as a baseline
- Fast inference; good for large-batch pre-screening
Weaknesses
- License explicitly bans commercial use — a hard veto
- 96×96 resolution; quality and detail trail everywhere
- The whole face region can degrade, not just the mouth
- No static-image-to-video support
3.4 LivePortrait Fastest — but not a lip-sync specialist Overall 7.8
It is not a “lip matching” model but a portrait animation driver: it transfers expressions, head pose, and eye motion from a driving video (or audio adapter) onto a still portrait. Its killer feature is the stitching module — the animated face is seamlessly pasted back onto the original torso, eliminating the classic “neck seam” tell. At ~12ms/frame, it is the only open-source option that makes real-time interactive applications viable.
The selection criterion is sharp: you want expression and head dynamics, not strict phoneme-level lip-sync. For VTubers, virtual-host expression driving, and photo animation, it is the best pick; for fully matching a long dubbing track, its lip accuracy does not make the first tier.
Strengths
- ~12ms/frame — untouchable speed, supports real-time apps
- Stitching module eliminates seam artifacts entirely
- 6–8GB VRAM; low barrier
- Expression, eye, and lip regions independently adjustable
- Handles real people, illustrations, and anime styles
Weaknesses
- Not an audio-native lip specialist; accuracy depends on the audio adapter
- Weight licensing is contradictory across sources — commercial risk must be self-verified
- Head-and-shoulders only; no full body or gestures
3.5 Tencent HunyuanVideo-Avatar Multi-character + emotion control Overall 7.9
All three of its design moves point at one goal: making avatars “perform,” not just “speak.” Character image injection replaces legacy additive conditioning, eliminating train/inference condition mismatch in exchange for cross-frame identity and larger motion. The Audio Emotion Module extracts mood from a single emotion reference image and transfers it to the target video. The Face-Aware Audio adapter uses face masks to isolate each character’s audio injection in latent space, enabling multiple characters speaking independently on screen — a rare capability on this list.
It already ships inside Tencent products (QQ Music AI Singer, Kugou storybooks, WeSing MV), so engineering maturity is credible. But two limits are hard: the 14-second audio input cap forces long content into sliced segments with seams that easily show in emotion and pose; and the license excludes the EU, UK, and South Korea — that one line decides project viability more than any spec.
Strengths
- Multi-character independent on-screen drive — a scarce capability
- Controllable emotion style (AEM transfers from a reference image)
- Full-body framings and multiple styles (realistic / cartoon / 3D / stylized)
- Proven in high-traffic Chinese products; stability backed by production
- Community quantization path lowers the bar to 10GB
Weaknesses
- License excludes EU / UK / South Korea — disqualifying for global launches
- Audio ≤14 s; long video requires slicing, and seams are hard to hide
- 24GB minimum, 96GB recommended — expensive hardware
- Output capped at 720p; not broadcast-grade
3.6 Alibaba Wan2.2-S2V-14B Minute-long video Strong Chinese performance Overall 7.9
Among open-source avatar models it owns one metric nobody else has: minute-long continuous generation. Most peers stop at “seconds”; Wan2.2-S2V uses hierarchical frame compression to drastically shrink historical-frame tokens, extending reference length from a few frames to 73 — so long videos stop degrading frame by frame and become usable for industrial scenarios like avatar livestreaming. In real tests, Chinese lip-sync lands on syllables, with feedback on stress and pauses; a few plosives stay muddy — “speaking,” not “reading a script.”
The cost is hardware: 14B parameters, officially 80GB per card, so consumer GPUs need quantization and sharding. It is also sensitive to audio quality — heavy background music, echo, or muddy vocals visibly drift the lips; denoise and dereverberate first.
Strengths
- The only open-source option with stable minute-long generation
- Good Chinese lip-sync and emotional feedback
- Supports real people, cartoons, animals; portrait to full body, any aspect
- Prompt control over motion trajectories and background
- Audio-visual sync keeps pace with full-body motion rhythm
Weaknesses
- Officially 80GB VRAM — the highest hardware bar here
- Output 480P–720P; no 4K delivery
- Slow generation; fast iteration needs smaller sizes or shorter clips
- Poor tolerance for low-quality audio; front-end cleanup is mandatory
3.7 EchoMimic V2 / V3 Half-body + gestures Apache-2.0 Overall 8.2
It fills the “half-body explainer” gap: most open models only nail head-and-shoulders, while the EchoMimic family brings gestures and half-body posture into generation, paired with a clean Apache-2.0 license — making it the best value open-source option for knowledge talking-heads and e-commerce explainers. V3 is only ~1.3B parameters, runs on 12GB, and batch-producing on a rented A40 pushes per-clip cost very low.
Its positioning is not precision champion — quality and stability trail LatentSync / MuseTalk — but at the intersection of “clean license + gestures + hardware friendly,” it is essentially the only solution.
Strengths
- Apache-2.0 with no additional commercial terms
- Half-body + gestures make explainer content more natural
- V3 is only ~1.3B params; 12GB suffices
- Predictable batch-production cost in the cloud
Weaknesses
- Quality and lip accuracy below the head of the field
- Long-video stability weaker than Wan2.2-S2V
- Gesture quality varies with reference image and audio rhythm
3.8 sync.so · sync-3 Commercial precision ceiling Overall 8.9
Judging purely on “make the mouth match,” it is the best commercial answer today, and its edge sits exactly where others cannot go: automatic occlusion detection (hands, microphones, and glasses over the mouth need no manual masks), extreme angles and partial faces, and the only capability to open a fully closed mouth — lipsync-2 and 2-pro both fail that case. It is also the only model accepting static image input, which amounts to built-in light talking-head generation.
Pricing is the steepest on the list: at 25fps, one minute of 4K output runs about $6.4–8; the free tier is 20 seconds with a watermark, so real use means the Creator plan ($19/mo, watermark off, 5-minute clips) plus metered overage. It explicitly does not support real-time, and it cannot rescue footage (other than stills) where “no natural speaking motion” exists.
Strengths
- Native 4K output, full-clip processing (not 2-second chunks)
- Auto occlusion detection; extreme angles and partial faces natively supported
- The only model that can open a fully closed mouth
- The only one with image input — image-to-talking-video
- Most complete developer toolchain (SDK / OpenAPI / plugins / MCP)
- Language-agnostic waveform analysis; 95+ languages without penalty
Weaknesses
- Most expensive: ~$6.4–8.0/min (4K)
- Dual billing (subscription + usage) needs threshold management
- Free tier is watermarked and capped at 20 s — evaluation only
- No real-time / livestream support
- Developer-oriented; no full non-technical interface
3.9 HeyGen End-to-end multilingual dubbing Overall 8.6
It sells not a model but a closed loop: upload the finished video → auto-transcribe → translate → clone the voice → lip-sync → download, with no third-party tools or self-assembled pipeline. The value peaks when “one video needs six languages” — under its public credit rules, a 90-second source plus six lip-synced dubs costs just 75 of 600 credits.
Know its boundary: HeyGen’s lip-sync exists to serve avatars and translation — reliable on single-speaker, frontal, clean footage, but it loses to specialists like sync-3 on complex camera moves, profiles, and multi-person occlusion. It also confirms an industry rule proven repeatedly: lip-sync doubles or triples per-minute cost. If no one’s face is on screen, audio-only dubbing plus accurate subtitles wins on value.
Strengths
- End-to-end loop; no pipeline to build yourself
- Low per-output-minute cost in class ($0.24 with lip-sync)
- 175+ languages and regional variants — the widest coverage
- Transparent credit system; costs are precisely forecastable
- Mature team collaboration, brand assets, enterprise governance
Weaknesses
- Lip-sync oriented to avatars and translation; poor at repairing complex live footage
- Free tier limited to 1 minute and watermarked
- Credits and quotas require manual math; easy to overestimate
- Extreme angles, occlusion, and multi-speaker scenes lose to specialists
3.10 Volcano Engine · OmniHuman 1.5 China’s value benchmark ¥1/second Overall 8.7
It is the absolute value benchmark on this list: ¥1/second means 1080P avatar video at ¥60/minute — far more transparent than domestic rivals’ “avatar annual fee + voice annual fee + overage” pricing (effectively ¥11.6–25.9 per short video), and much cheaper than sync-3’s $6.4–8/minute.
The trade-offs are written in the same documentation: audio must be under 60 seconds, with quality degrading past 15 seconds, and concurrency is capped at 1 — batch production queues up. The docs also flag two traps: faces too small in frame intermittently produce “no lip movement” (the character stops talking), and structural stability decays past 15 seconds, with re-entering characters showing identity drift. These are not dirt — they are constraints to design around.
Strengths
- ¥1/s — China’s avatar value benchmark
- 1080P output with lips + emotion + gestures in one
- Multi-character performance and camera control
- Clear domestic onboarding and compliance path (Jimeng console)
- Per-second billing — no avatar-year fee or overage math traps
Weaknesses
- Audio <60 s, degrading past 15 s; long content must be sliced
- Concurrency capped at 1 — batch production queues
- Very wide shots / small faces intermittently stop lip movement
- Stability decays past 15 s; re-entering characters drift
3.11 Hosted Platforms and China’s Avatar SaaS Matrix
Beyond self-hosting and direct vendor APIs, two middle routes are common.
| Route | Typical offerings | Price reference | Best for |
|---|---|---|---|
| Model aggregators | fal.ai, Replicate, muapi etc. | Kling LipSync ~$0.014/s; Veed ~$0.52/min; Sync ~$0.91/min; LatentSync ~$0.26/clip (≤40 s); OmniHuman 1.5 ~$0.045–0.06/s | No GPU to babysit, want to A/B many models, spiky usage |
| China avatar SaaS | Baidu XiLing, Tencent Zhiying, SenseTime Ruying, Silicon Intelligence, iFlytek Zhizuo, ShanJian | Avatar annual fee ¥3,598–7,999 (500–1500 min incl.) + voice fee; effectively ¥3.7–25.9 per short video | Gov/enterprise and compliance scenarios, on-prem needs, livestream commerce |
Beware the conversion trap: Chinese SaaS vendors bill on a three-layer “avatar × voice × minute quota” model, and voice quota almost always runs out before avatar quota — the same short video looks cheaper on an E-tier avatar, but that tier’s 100-minute quota and lower quality, with ¥5/min overage, ends up more expensive.
4. Cross-Comparison Matrix
4.1 Method Decides Everything: Four Technical Routes
| Method | Representatives | Mechanism | Speed | Detail ceiling | Input |
|---|---|---|---|---|---|
| GAN | Wav2Lip | Generator + sync discriminator generate the mouth directly | Fastest | Low (96×96) | Video |
| Latent inpainting | MuseTalk | Single-step inpainting of the lower face; no iteration | Real-time (30fps+) | Medium (256×256) | Video |
| Latent diffusion | LatentSync, HighSync | 20–50 denoising iterations + audio conditioning | Slow | High (512×512) | Video |
| 3DMM + rendering | SadTalker | Audio predicts 3D coefficients, then the face is rendered | Medium | Medium | Single image |
Read this table against the model list and many “whys” answer themselves: MuseTalk is fast but stuck at 256 because it paints in one step; LatentSync is slow but sharp because diffusion iterates; only the 3DMM family eats static images because it builds a 3D representation first instead of editing pixels.
4.2 Commercial Price Normalization (per output minute)
| Service / model | Listed metering | Per output minute | Relative cost |
|---|---|---|---|
| Kling LipSync (fal.ai) | $0.014/s | ~$0.84/min | Lowest |
| Volcano Engine OmniHuman 1.5 | ¥1/s (China) | ¥60/min (~$8.4) | Best in China |
| HeyGen (audio-only dubbing) | 2 credits/min | ~$0.10/min | Lowest (no lip-sync) |
| HeyGen (lip-synced translation) | 5 credits/min | ~$0.24/min | Low |
| Vozo AI (with lip-sync) | Credits-based | ~$0.44–0.60 + lip add-on $1.5–2.0 | Mid |
| Rask AI | $150/mo / 100 min | ~$1.50/min | Mid |
| Synthesia | Credits-based | Up to ~$2.90–2.97/min | High |
| sync.so lipsync-2 | $0.04–0.05/s | ~$2.40–3.00/min | Mid-high |
| sync.so lipsync-2-pro | $0.067–0.083/s | ~$4.0–5.0/min | High |
| sync.so sync-3 (4K) | $0.107–0.133/s | ~$6.4–8.0/min | Highest |
| Self-hosted LatentSync / MuseTalk | Electricity + GPU depreciation only | Depends on utilization | Lowest at scale |
One sentence summarizes the table: lip-sync multiplies per-minute cost by 2–3×, and 4K doubles it again. Industry testing confirms it — lip-sync pays off when the subject’s face is on camera; on voiceover-plus-B-roll content, plain dubbing with accurate subtitles beats lip-synced dubbing.
4.3 Open-Source Academic Benchmarks (HDTF and public sets)
| Solution | FID ↓ (quality) | LSE-C ↑ (lips) | CSIM ↑ (identity) | Silence test ↑ (higher is better) |
|---|---|---|---|---|
| Wav2Lip | 14.912 | 7.63 | 85.2 | 0.84 |
| VideoReTalking | High | — | — | 0.82 |
| MuseTalk 1.5 | 8.759 (paper: 6.52) | 6.89 | 86.2 | 0.68 |
| LatentSync 1.6 | 8.518 | 8.05 | 85.9 | 0.81 |
| Diff2Lip | 12.079 | 7.14 | 86.9 | 0.78 |
| HighSync | 7.36 (HDTF) | 7.72 | 0.86 | 0.93 |
FID varies considerably across sources due to implementation and preprocessing differences (MuseTalk self-reports 6.52; third-party boards list 8.759); cross-table numbers are not directly comparable. LSE-C units also differ (some 0–10, some percentages).
5. Real-World Data: What the Marketing Skips
5.1 The Four Traps of LSE-C / LSE-D
These two metrics come from the Wav2Lip paper and are the most-cited lip-sync metrics anywhere. Treating them as procurement evidence runs into four traps:
- Low correlation with human judgment. Papers reviewing lip-sync evaluation frameworks state outright that LSE-C and LSE-D correlate “very limitedly” with subjective human scores — some methods even score far above ground truth on them; the metrics are distorted enough to be beaten.
- SyncNet is not translation-invariant. It is sensitive to crop windows, face position, brightness, image quality — even codec choice. The eye sees no difference; the score has already moved.
- Circular reasoning. When a model is trained to push SyncNet’s output toward 1 (the lip-sync loss) and then graded by the same SyncNet’s LSE-C / LSE-D, the reported “best” is merely its training objective, not an independent test.
- Better features exist. Follow-up work uses AV-HuBERT audio-visual features with cosine similarity (AVSu), whose representations are more stable and shift-robust than SyncNet’s.
5.2 The Silence Test: The Best One-Shot Veto
The procedure is trivial: feed the model silence and watch the mouth. If it keeps moving, the model is inferring lip trajectories from visual context rather than truly conditioning on audio — the notorious data-leakage problem of diffusion lip models (LatentSync’s SyncNet supervision was designed precisely to fix it). From public comparisons:
| Solution | Silence-test score | Reading |
|---|---|---|
| HighSync | 0.93 | Heavy normalization and masked attention; most reliable silence behavior |
| Wav2Lip | 0.84 | Weak but manageable |
| VideoReTalking | 0.82 | Same tier |
| LatentSync | 0.81 | Among the better diffusion models |
| Diff2Lip | 0.78 | Clear mouth movement during silence |
| MuseTalk | 0.68 | Worst — single-step generation lacks true temporal audio modeling |
This test has enormous practical value: it predicts “is the final cut watchable” better than any LSE number. For content with pauses, silence, and conversational gaps (interviews, courses, meeting edits), always run this first.
5.3 Where Human Evaluation Diverges from Automatic Metrics
The HighSync paper’s human study surfaced a highly revealing split:
| Solution | Quality (human, /5) | Lip-sync (human, /5) |
|---|---|---|
| Ground Truth | 4.78 | 4.35 |
| MuseTalk | 4.34 (highest) | 3.14 (lowest) |
| HighSync | 4.28 | 4.01 (highest) |
| LatentSync | — | 3.68 |
| Wav2Lip | 3.78 | — |
| Diff2Lip | 2.15 (lowest) | — |
Three takeaways, all practical: ① quality and lip-sync are independent curves — MuseTalk’s adversarial refinement makes the prettiest frames and the worst mouths; ② real video is not an unreachable ceiling — HighSync’s sync score (4.01) and human quality (4.28) approach ground truth (4.35 / 4.78) while remaining perceptibly behind; ③ low-resolution pixel-space diffusion shows under the human eye — Diff2Lip’s 2.15 was the lowest score of the entire study.
5.4 What Happens on Real Footage
- Audio quality is variable #1. Vendor docs repeat it: heavy background music, reverb, or muddy vocals visibly drift the lips. Wan2.2-S2V’s practical advice is “denoise and dereverberate before feeding.”
- Very wide shots are a minefield. Volcano Engine explicitly warns that tiny faces intermittently produce no lip movement; most models share this — too few mouth pixels to discriminate.
- Duration is a hard boundary. Volcano audio must be <60 s (≤15 s recommended), Hunyuan audio ≤14 s, MOVA ≤8 s per clip, while sync.so’s lipsync-2 family processes in independent 2-second chunks — check the seams between chunks.
- 15 seconds is the universal stability knee. Volcano and Kling-family docs alike note structural stability and consistency decaying after 15 seconds — not one vendor’s flaw but the current state of temporal modeling.
6. Scoring and Head-to-Head Evaluation
The previous chapters are the material; this one is the verdict. A disclaimer first: the scores below are a synthesized evaluation, not official benchmarks. They draw on vendor documentation and pricing pages, model cards and papers, repository licenses and READMEs, and third-party production tests and human-eval studies — closer to a “procurement-view comparability score” than an academic benchmark.
6.1 Dimensions and Weights
| Dimension | Weight | What it measures |
|---|---|---|
| Lip accuracy | 25% | Phoneme-level alignment, silence behavior, opening closed mouths, stability at extreme angles |
| Visual fidelity | 20% | Output resolution, teeth/lip detail, identity preservation, whole-face degradation |
| Licensing & commercial use | 15% | Commercial permission, regional exclusions, revenue/MAU gates, weight-inheritance risk |
| Hardware & speed | 15% | VRAM barrier, real-time or not, throughput, local/on-prem feasibility |
| Ease & ecosystem | 10% | Hosted endpoints/APIs/plugins, self-build burden, community activity, documentation |
| Scenario coverage | 10% | Re-dubbing / single-image generation / multi-speaker / long video / half-body gestures |
| Cost | 5% | Blended per-output-minute cost (subscriptions, quotas, overages) |
6.2 Overall Scoreboard
| Rank | Solution | Overall | Stars | One-line verdict |
|---|---|---|---|---|
| 1 | sync.so sync-3 |
8.9
|
★★★★★ | Only 4K + auto occlusion; the price is being the most expensive of all |
| 2 | Volcano Engine OmniHuman 1.5 |
8.7
|
★★★★☆ | 1080P avatars at ¥1/s — no domestic rival |
| 3 | HeyGen |
8.6
|
★★★★☆ | Not the best lips — the least friction multilingual loop |
| 4 | LatentSync 1.6 |
8.4
|
★★★★☆ | Open-source re-dubbing’s quality + license double win; loses on 18GB VRAM |
| 5 | MuseTalk 1.5 |
8.3
|
★★★★☆ | MIT + truly real-time; the 0.68 silence score is the hard flaw |
| 6 | EchoMimic V2 / V3 |
8.2
|
★★★★☆ | Clean license + half-body gestures + 12GB — the only solution at that intersection |
| 7 | Tencent HunyuanVideo-Avatar |
7.9
|
★★★☆☆ | Strong multi-character + emotion, held down by a three-country license exclusion |
| 8 | Alibaba Wan2.2-S2V-14B |
7.9
|
★★★☆☆ | The only minute-long video, but the 80GB bar turns most people away |
| 9 | LivePortrait |
7.8
|
★★★☆☆ | Untouchable at 12ms/frame — but it’s expression transfer, not lip refinement |
| 10 | Wav2Lip |
6.3
|
★★☆☆☆ | Defined the field — and is locked out of commerce by its own license |
6.3 Seven-Dimension Scoreboard
| Solution | Lips 25% |
Quality 20% |
License 15% |
Hardware 15% |
Ease 10% |
Coverage 10% |
Cost 5% |
Overall |
|---|---|---|---|---|---|---|---|---|
| sync-3 | 9.5 | 9.5 | 8.0 | 9.5 | 9.0 | 9.2 | 4.0 | 8.9 |
| OmniHuman 1.5 | 9.0 | 9.2 | 8.5 | 9.5 | 8.8 | 8.0 | 5.5 | 8.7 |
| HeyGen | 8.3 | 8.8 | 8.5 | 9.5 | 9.5 | 7.0 | 7.5 | 8.6 |
| LatentSync 1.6 | 9.3 | 9.0 | 9.5 | 6.0 | 7.0 | 7.5 | 9.5 | 8.4 |
| MuseTalk 1.5 | 8.0 | 8.2 | 9.5 | 9.0 | 7.5 | 7.0 | 9.5 | 8.3 |
| EchoMimic V2/V3 | 8.2 | 8.5 | 9.5 | 7.5 | 6.5 | 8.0 | 9.0 | 8.2 |
| HunyuanVideo-Avatar | 8.6 | 8.8 | 6.5 | 7.0 | 7.0 | 8.5 | 9.0 | 7.9 |
| Wan2.2-S2V-14B | 8.5 | 9.0 | 9.0 | 5.0 | 6.5 | 8.0 | 9.0 | 7.9 |
| LivePortrait | 7.0 | 8.8 | 6.0 | 9.8 | 8.0 | 6.0 | 9.5 | 7.8 |
| Wav2Lip | 7.5 | 4.0 | 2.0 | 9.5 | 8.5 | 5.0 | 10.0 | 6.3 |
The column to stare at is “License”: Wav2Lip 2.0, LivePortrait 6.0, Hunyuan 6.5 — three numbers representing three distinct traps: explicit commercial ban, disputed weight licensing, and regional exclusions. They have nothing to do with lip accuracy, yet they decide project survival more than any other column.
6.4 Category Champions
Native 4K, auto occlusion, the only one that opens closed mouths; extreme angles and partial faces natively handled.
Highest human-judged quality at 4.34 (ground truth: 4.78); double first on FID and CSIM over HDTF.
MIT and Apache-2.0 — no regional exclusions, no revenue gates. The only three open-source options that go straight into commercial projects.
~12ms/frame on 6–8GB VRAM — fast enough for real-time interaction on a 4090; the “lightest” option on this list.
Head-and-shoulders to full body, multi-character independent drive, emotion transfer — the most capabilities per card.
The only open-source minute-long stable generation; Chinese lips land on syllables with audible stress and pause feedback.
6.5 Head-to-Head Verdicts (Four Matchups)
① Re-dubbing real footage: sync-3 vs LatentSync 1.6
Quality and hard cases: sync-3 wins — 4K, auto occlusion, extreme angles. Cost: LatentSync wins big — self-hosting costs electricity against $6.4–8/min, a 20×+ gap. LatentSync’s weakness is 18GB VRAM plus ops burden; sync-3’s weakness is budget and real-time.
② Single-image avatar: OmniHuman 1.5 vs HunyuanVideo-Avatar vs Wan2.2-S2V
For stable output and domestic compliance, OmniHuman 1.5 wins (¥1/s, 1080P); for multi-character and emotion control, HunyuanVideo-Avatar wins; for long video and Chinese lips, Wan2.2-S2V wins. Shared weakness: duration — 60 s, 14 s, and 8 s caps all force slicing and stitching.
③ Real-time / livestream: MuseTalk 1.5 vs LivePortrait
To “speak in real time with the audio,” MuseTalk wins — audio-native at 30fps+. To “transfer expressions and head motion from a driving video,” LivePortrait wins — 12ms/frame with clean stitching. Don’t swap them: LivePortrait refining long dubbing lips, or MuseTalk doing expression transfer, both yield second-best results.
④ The open-source license matchup: the three clean ones vs the traps
MuseTalk (MIT), LatentSync (Apache-2.0), EchoMimic (Apache-2.0) go straight into commercial projects. Wav2Lip explicitly bans commercial use (LRS2 training), VideoReTalking is CC BY-NC-SA, SadTalker’s weights inherit upstream restrictions, Sonic requires Tencent Cloud for commercial use. Prototyping on Wav2Lip is fine — replace it before shipping.
7. Engineering Practice: The Six Factors That Decide Success
| Factor | Why it matters | How to handle it |
|---|---|---|
| ① Audio quality | The #1 source of lip drift; most models have poor tolerance for low-quality audio | Denoise, dereverberate, normalize to 16kHz; ensure a single dominant voice |
| ② Occlusion | Hands, mics, or glasses over the mouth break nearly every model | Only sync.so’s sync-3 detects occlusion automatically; the rest need manual masks or reframed shots |
| ③ Starting mouth pose | If the source’s first frame has a closed mouth, most models cannot open it | lipsync-2 / 2-pro can’t; only sync-3 can — or trim the closed opening frames |
| ④ Face scale and angle | Tiny faces produce “no lip movement”; extreme angles and profiles degrade broadly | Prefer medium close-up frontal shots; for wide shots, reshoot or switch to the avatar-generation track |
| ⑤ Multiple people in frame | One wrong attribution ruins the whole clip | Use Active Speaker Detection (sync.so Creator tier and up) or HunyuanVideo-Avatar’s FAA / MultiTalk |
| ⑥ Duration and seams | 15 s is the universal stability knee; chunked processing shows at seams | Cut by scene, not by seconds; align seams to semantic boundaries; for long videos use Wan2.2-S2V or full-clip sync-3 |
8. Recommendations by Scenario
Multilingual dubbing for finished videos
Choose sync-3 for 4K and hard shots; self-host LatentSync when GPUs and volume exist — a 20×+ cost difference.
Single-image talking avatar
1080P at ¥1/s with a clear domestic compliance path. Mind the <60 s audio cap, concurrency 1, and avoid ultra-wide shots.
Multi-speaker conversations
Tencent’s FAA isolates each character’s audio injection with face masks for independent multi-character drive; MultiTalk is lighter but caps at ~15 s per clip.
Live real-time lip-sync
The only true real-time option (30fps+), MIT licensed, 8GB VRAM. Mouths move through silence — tolerable in a livestream.
Finished ads, extreme angles
Native 4K, auto occlusion, partial faces and extreme angles — the only commercial option that covers “cannot reshoot” cases.
Chinese short video, cost-sensitive
Self-host Wan for good Chinese lips and minute-long output; skip the GPU with Volcano’s per-second billing.
Half-body explainers, e-commerce
Apache-2.0 + gesture generation + 12GB VRAM; batch-producing on a rented A40 keeps per-clip cost very low.
VTubers, expression driving
12ms/frame, seamless stitching, independently adjustable expression/eyes/lips. When you want motion, not phoneme-level precision, it’s the pick.
Corporate training, multilingual courses
175+ language closed loop, $0.24/min with lip-sync, enterprise and brand governance — best for “steady output at volume.”
Data must stay on-prem
Apache-2.0 and MIT dual licenses, fully local pipeline; air-gapped deployment is rare on the commercial side.
Just validating the effect first
Kling LipSync ~$0.014/s, LatentSync ~$0.26/clip (≤40 s) — compare many models without owning a GPU.
Gov/enterprise and compliance first
On-prem deployment and SLAs available; but compute the three-layer “avatar × voice × minute quota” pricing — the voice quota overruns first.
Decision Framework: Four Steps to a Shortlist
Step 1 — classify your task. Do you have “an existing video that needs a new mouth” or “only a single image”? The former: LatentSync / MuseTalk / sync-3; the latter: OmniHuman / HunyuanVideo-Avatar / Wan2.2-S2V. Getting this box wrong wastes everything after it.
Step 2 — subtract by license. For commercial use, strike Wav2Lip, VideoReTalking, SadTalker, and Sonic first, then check regional clauses — Hunyuan excluding the EU/UK/South Korea is an outright veto for global launches.
Step 3 — subtract by footage conditions. Occlusion, extreme angles, closed starting mouth → only sync-3; real-time needed → only MuseTalk / LivePortrait; minute-long video needed → only Wan2.2-S2V.
Step 4 — find the cost break-even. At 4K’s $6.4–8/min, above roughly 30 minutes per month with stable GPUs, self-hosting’s electricity + ops starts winning; below that, hosted convenience is worth more.
9. Pitfall Checklist and Compliance
9.1 License Traps, Line by Line
| Model | Apparent license | Actual status |
|---|---|---|
| Wav2Lip | Publicly downloadable repo | Explicitly non-commercial (trained on LRS2); third-party forks redistributing it do not transfer authorization |
| VideoReTalking | Academic open source | CC BY-NC-SA 4.0 — non-commercial |
| SadTalker | Code Apache-2.0 | Weights fine-tuned from SD 1.5 (OpenRAIL-M field restrictions) + PIRenderer (research-only); inherits upstream constraints |
| MuseTalk | MIT | Model is commercially usable; bundled test assets are research-only — don’t ship them with deliverables |
| Sonic (Tencent) | Open and downloadable | Non-commercial only; README states commercial use must go through Tencent Cloud’s video creation model API |
| LivePortrait | Code MIT | Weight licensing is contradictory across sources (MIT / non-commercial research); verify the repo’s LICENSE text before commercial use |
| Duix.Avatar | Custom license | Free under 100k users and $10M revenue; commercial license beyond |
| HunyuanVideo-Avatar | Open weights | Tencent Hunyuan Community License, excluding EU / UK / South Korea; separate agreement above 100M MAU |
| LatentSync | Apache-2.0 | Commercial OK; note that likeness and voice rights of the person on screen are a separate layer |
9.2 Three Compliance Lines
- Labeling obligations (China). The Provisions on Deep Synthesis in Internet Information Services require labeling deep-synthesis content; re-mouthed videos are a textbook case — implement labeling and platform filing before publishing.
- Transparency (EU). The AI Act imposes disclosure duties on deepfake content; international releases need machine-readable or visible labeling.
- Likeness and voice authorization. Entirely independent of licensing: a model’s license permitting commercial use does not mean you may drive a real person’s face and voice. Explicit consent is required, and celebrity likenesses and voices should not be touched at all.
One practical compliance rule from industry testing: subtitle and translation accuracy moves perceived quality more than lip-sync does. Testers rarely notice missing lip-sync, but they notice a mistranslated product name instantly — so budget translation quality above lip fidelity.
Key Findings
1. This is four tracks, not one market. Re-dubbing (fix the mouth), avatar generation (create a performance), joint generation (audio and video at once), and end-to-end dubbing (the whole pipeline) differ completely in inputs and outputs — mixing them yields second-best results. Step one is always classifying your own task.
2. Method and license matter more than parameters. GAN is fast but only 96×96; single-step inpainting is real-time but stuck at 256; diffusion is sharp but iterative and 18GB-hungry; 3DMM is the only one that eats still images. And what eliminates most candidates is usually the license, not the technology.
3. Don’t buy based on LSE-C / LSE-D. They correlate poorly with human judgment, are sensitive to cropping and brightness, can be “beaten” above ground truth, and carry the circularity of training and scoring on the same SyncNet. The silence test, human blind evaluation, and your own footage are better judges.
4. The silence test is the highest-value veto. MuseTalk’s 0.68 versus HighSync’s 0.93 is the difference between “truly listening” and “inferring from pixels.” Any content with pauses or silence should run this test before any other metric is discussed.
5. The open-source / commercial boundary sits at ~30 minutes per month. 4K commercial runs $6.4–8/min; self-hosting costs electricity but demands 18GB VRAM and maintenance. Small volume favors hosted; large volume with someone to babysit GPUs favors self-hosting.
6. China’s answer is different. Volcano Engine delivers 1080P avatars at ¥1/s with concurrency 1 and a 60-second audio cap — ideal for batch short-video production, a poor fit for long-form or high concurrency. Selecting domestic solutions by Western leaderboards selects wrongly.
7. The last step is always your own footage. Every public number was earned on specific datasets; real business footage challenges you on audio quality, occlusion, angle, and duration simultaneously. Run one clip of your own material before trusting any ranking.
Lip-sync solves the mouth. Who keeps the voice?
Once the video is localized, the voice has to sound like the original speaker. DeepVideo — AI Video Translation with Voice Clone and TTS in 30+ languages, running locally on your machine — no cloud upload.
Free tier: 18 minutes total + 2 minutes daily · Local processing, no cloud upload
© 2026 DeepForgeHub Research. Sources: ByteDance LatentSync paper and model card (arXiv:2412.09262), Tencent MuseTalk paper (arXiv:2410.10122), Tencent Hunyuan HunyuanVideo-Avatar technical report and repository, Alibaba Tongyi Wan2.2-S2V official releases, Ant Group EchoMimic V2/V3 repositories, OmniHuman-1 and HighSync papers, sync.so pricing and model documentation, HeyGen / D-ID / Synthesia / Rask AI public pricing pages, Volcano Engine Jimeng OmniHuman 1.5 billing documentation, Baidu XiLing / Tencent Zhiying / SenseTime Ruying / Silicon Intelligence public quotes, and third-party human-evaluation and silence-test studies. Prices are public list prices; metric definitions are annotated in the text; as of September 2026.

