Voice-Pro - AI Voice Cloning and Dubbing Tool
Voice-Pro is an AI tool for voice cloning and dubbing. It supports multiple languages, integrates with Gradio, and can process YouTube videos for easy content creation.

Introduction
Voice-Pro is an all-in-one AI voice processing toolbox that covers the entire workflow from speech recognition to dubbing. In the field of voice processing, creators typically need to switch between multiple tools: downloading videos, separating vocals, transcribing speech, translating subtitles, and then synthesizing dubbed audio. Voice-Pro consolidates all of these steps into a single, ready-to-use web application built on Gradio. It supports YouTube video downloading, vocal separation, speech recognition, translation, subtitle generation, and multilingual dubbing, providing content creators, researchers, and cross-border teams with a complete AI voice processing pipeline.
Key Features
- Top-tier speech recognition: Integrates Whisper, Faster-Whisper, and Whisper-Timestamped engines, delivering high-accuracy transcription with word-level timestamps and strong performance in Chinese, English, and other languages.
- Zero-shot voice cloning: Built-in F5-TTS, E2-TTS, and CosyVoice (including Fun-CosyVoice3 supporting 9 languages such as Korean) allow you to clone any voice without training samples.
- Multilingual text-to-speech: Powered by Edge-TTS and kokoro, with optional Azure TTS integration, to generate natural-sounding dubbing in multiple languages with a single click.
- YouTube video processing: An integrated downloader (yt-dlp) lets you pull videos directly for subtitle generation, translation, and re-dubbing to create multilingual content.
- Professional vocal separation: Incorporates Demucs and MDX-Net models for accompaniment/vocal splitting, suitable for karaoke, podcast editing, and mixing scenarios.
Highlights
- Open-source under the LGPL license, with all models and processing running locally on Windows with CUDA acceleration support.
- Privacy-first design: user data never needs to be uploaded to the cloud, ensuring full control and security.
- Near ElevenLabs-level cloning and dubbing capability without monthly subscription fees — deploy once and use long-term.
- Pipeline-style task design significantly boosts efficiency for batch processing scenarios.
- Over 11,000 GitHub stars with an active community and continuous updates.
Who It's For
Content creators can use Voice-Pro to batch-generate multilingual subtitles and dubbing for quick global reach; podcasters and video editors can streamline material processing with vocal separation and transcription; language learners can immerse themselves through video translation and bilingual subtitles; and developers can build on its modular architecture to create their own voice applications. If you are looking for a free, locally deployable AI voice workflow tool, visit the Voice-Pro GitHub repository today and experience the one-stop workflow from video to multilingual dubbing.





