DeepVideo 2.0 Beta: Tauri 2 + Rust Rebuild, 77× Faster Startup
DeepVideo 2.0 Beta coming soon - original video next to AI translated video with lip movements preserved

DeepVideo 2.0 Beta: A Complete Rebuild of Video Translation & Dubbing

DEEPVIDEO 2.0 BETA · COMING SOON

A Complete Rebuild — A New Era of Video Translation & Dubbing

Tauri 2 + Rust native rebuild · 77× faster startup · 100% offline processing

Release date: October 1, 2026 | Platforms: macOS (Apple Silicon / Intel) · Windows

Summary: DeepVideo 2.0 is a complete rebuild from the ground up. We ditched the old Python + FastAPI architecture in favor of Tauri 2 + Rust to create a truly native desktop app. The change isn’t just numbers on paper — it’s speed and fluidity you can feel every time you open the app and process every video.

  • 77× faster startup — navigation + skeleton screen visible within 73ms
  • Dramatically lower memory usage — streaming voice separation cuts memory for 45-minute audio from 3.4GB to 70MB
  • 100% offline processing — your video data never leaves your computer
  • Consistent cross-platform experience — macOS and Windows fully feature-aligned

Brand-New Features

DeepVideo 2.0 project list — brand-new feature modules
DeepVideo 2.0 project list: Expert Video Translation · Video Download · Image Watermark Removal · Video Watermark Removal

🎓 Expert Video Translation Pro Workflow

A full-pipeline translation workflow for professional users: a 10-step pipeline with advanced translation options — set translation style, target industry, and glossary; native multi-speaker mode and emotion recognition, plus pause-after-translation review so you can review before synthesis continues.

⬇️ Video Download Free to UseBuilt-in yt-dlp

Built-in yt-dlp tool for one-click downloads from major platforms, with the option to create a translation project right after downloading — download and translation closed-loop in one app, no third-party tools needed. This feature is free to use — no subscription required.

🖼️ Image Watermark Removal Free to UseLaMa Inpainting

AI automatically detects watermark locations and removes them with the LaMa inpainting algorithm, with batch processing support — import many images at once, export in one go. This feature is free to use — no subscription required.

🎬 Video Watermark Removal Free to UseFrame-by-Frame AI Erasing

Smart detection of hardcoded subtitles and watermarks with frame-by-frame AI erasing: the new GrowPolicy reduces accidental damage by 61–87%, validated on a corpus of 10 videos / 118 frames. This feature is free to use — no subscription required.

Core Upgrades

🎙️ SenseVoice Speech Recognition Recognition Engine

The brand-new SenseVoice engine is now live. For casual Chinese speech, noisy environments, phone recordings and more, its recognition accuracy far exceeds Whisper:

  • Completely fixes Whisper’s homophone errors — cases like mis-hearing that no bigger model could ever fix are gone for good
  • Cantonese support — where Whisper forces Cantonese into Mandarin, SenseVoice outputs authentic Cantonese text
  • Smart routing: Chinese, Cantonese and Japanese automatically go to SenseVoice; other languages go to Whisper
  • Word-level timestamps for more precise subtitle alignment
Homophone errors: completely fixed (Whisper’s chronic issue)
Cantonese: authentic Cantonese text, no more forced Mandarin
Smart routing: Chinese/Cantonese/Japanese → SenseVoice, others → Whisper
Timestamps: word-level alignment; 196s audio recognized in ~26s (4 threads)

🗣️ Chatterbox Voice Cloning

The new Chatterbox voice cloning engine is fully integrated into the production pipeline:

Reference audio: clones a voice from just 4.5–6 seconds
Output quality: 40ms tail trimming for cleaner output
Universal controls: per-segment speed control + loudness normalization (ITU-R BS.1770)

🎶 RoFormer Vocal Separation New Default Engine

RoFormer replaces Demucs as the default vocal separation engine, validated by A/B blind testing on 19 audio clips: RoFormer 14 wins / 4 losses vs Demucs 11 wins / 7 losses.

  • Fixes female vocals being cut off and voiceovers being swallowed
  • Significantly less residual vocals in background music
  • Streaming processing cuts memory for 45-minute audio from 3.4GB to 70MB (48×)

👥 358 System Voice Characters Voice Library

A brand-new system voice library covering 10 languages (Chinese, English, Japanese, Korean, German, French, Spanish, Italian, Portuguese, Russian) and 8 categories (characters, narration, business, education, entertainment, news, stories, animals), with a built-in smart assignment algorithm that matches the most suitable voice based on speaker gender.

⚡ One-Click Translation Minimal Workflow

The brand-new one-click translation mode compresses the entire pipeline into a single step:

  1. Select a video, click translate
  2. Automatically: audio extraction → vocal separation → speech recognition → translation → speech synthesis → muxing
  3. After translation, go straight to the preview page with original and translated videos side by side
  4. Need fine-tuning? Jump into the subtitle editor with one click

Want to see what others have created with DeepVideo? Browse the Works showcase; new users can get started with the DeepVideo User Guide.

Experience Improvements

✏️ Editor Upgrades Editor

  • Per-segment TTS preview — no need to regenerate the whole audio track; click and listen per segment
  • Smart edit journal — only writes back changed settings, never overwrites unmodified ones
  • SRT import/export — supports SRT files with speaker info
  • LLM source-subtitle correction — automatically fixes speech recognition errors (original_corrected.srt)
  • Font size control — adjust subtitle styling in real time

🧽 Hard-Subtitle Removal Improvements Erasing Algorithm

Major improvements to the hard-subtitle erasing algorithm: the new GrowPolicy reduces accidental damage by 61–87%; the mask retention cap leaves no residue after subtitles leave the erasing zone; exit frames no longer freeze the background. Validated on a corpus of 10 videos / 118 frames plus 363-frame end-to-end tests.

👥 Speaker Diarization Improvements Diarization

  • Dual-metric merging — weak clusters use group cosine ≥ 0.35, strong clusters use pairwise centroid cosine ≥ 0.45, fixing different speakers being fused together
  • Iterative splitting at every speaker boundary, not just the first
  • Ignores sub-1-second bursts, eliminating ghost speakers

🍎 Mac App Store in Preparation App Store

DeepVideo 2.0 is being prepared for Mac App Store release: fully sandbox-compatible builds, security-scoped bookmarks persisting file access across restarts, legacy data migration shrunk from ~32GB to ~4MB, and the privacy manifest (PrivacyInfo.xcprivacy) already included.

Performance at a Glance

Metric DeepVideo 1.x DeepVideo 2.0 Improvement
Startup time ~5.6s ~73ms 77×
Vocal separation memory (45min audio) 3.4 GB 70 MB 48×
Project list DOM nodes Full rendering Paginated, 30 per page −86%
FFmpeg audio loading 1,432 MB/h 220 MB/h 6.5×
SenseVoice recognition (196s audio) — ~26s (4 threads) —

For the full changelog, see the Release Notes.

Supported Languages

UI languages: 20 | Speech recognition: auto-detected, supporting Chinese, Cantonese, English, Japanese, Korean, German, French, Spanish and more | Speech synthesis: 10 languages

Final Words

From 1.0 to 2.0, we did more than swap the tech stack. We rethought how “video translation” should be done — less waiting, lower resource usage, more accurate recognition, more natural voices. Our thanks to every beta tester. If you run into any issues, feel free to contact us. See you on October 1.

Further Reading: DeepForgeHub Research Reports

More Research Reports

See You on October 1

DeepVideo 2.0 Beta launches on October 1, 2026. The Mac App Store version is coming soon.

Download from the Official Site

Free quota: 18 minutes + 2 minutes daily | 100% local processing, never uploaded to the cloud

Newsletter Updates

Enter your email address below and subscribe to our newsletter

Leave a Reply

Your email address will not be published. Required fields are marked *