Voice-Pro - 开源AI语音识别、翻译与多语言配音工具

Voice-Pro - 开源AI语音识别、翻译与多语言配音工具

Voice-Pro - 开源AI语音识别、翻译与多语言配音工具 截图

Voice-Pro:一站式AI语音处理工具箱,从识别到配音全流程覆盖

在语音处理领域,创作者通常需要在多个工具之间来回切换:下载视频、分离人声、转写文字、翻译字幕、再合成配音。Voice-Pro 把这些环节全部整合进一个开箱即用的 Web 应用——基于 Gradio 构建,支持 YouTube 视频下载、人声分离、语音识别、翻译、字幕生成与多语言配音,为内容创作者、研究者和跨国团队提供了一条完整的 AI 语音处理流水线。

核心功能:覆盖语音处理全链路的五大能力

  • 顶级语音识别:集成 Whisper、Faster-Whisper、Whisper-Timestamped 三种引擎,支持高精度转录与逐词时间戳输出,中英文及多语种识别表现优异。
  • 零样本语音克隆:内置 F5-TTS、E2-TTS、CosyVoice(含支持韩语等 9 种语言的 Fun-CosyVoice3),无需训练样本即可克隆任意音色。
  • 多语言文本转语音:基于 Edge-TTS 与 kokoro,可选接入 Azure TTS,一键生成自然流畅的多语种配音。
  • YouTube 视频处理:内置下载器(yt-dlp),可直接拉取视频进行字幕生成、翻译与二次配音,打造多语言内容。
  • 专业人声分离:集成 Demucs 与 MDX-Net 模型,支持伴奏/人声分离,适配卡拉 OK、播客剪辑与混音场景。

特色优势:开源、本地运行、隐私安全

Voice-Pro 采用 LGPL 开源协议,全部模型与处理均在本地运行(Windows 平台,支持 CUDA 加速),用户数据无需上传云端,隐私安全可控。相比商业配音 SaaS,它提供了近乎 ElevenLabs 级别的克隆与配音能力,却无需按月订阅,一次部署长期使用。对于批量处理场景,其流水线式的任务设计也能显著提升效率。

适用人群:谁最适合使用 Voice-Pro?

内容创作者可用它批量生成多语言字幕与配音,快速出海;播客与视频剪辑师借助人声分离与转录功能高效处理素材;语言学习者可以通过视频翻译与双语字幕沉浸式学习;开发者则可基于其模块化架构二次集成,构建自己的语音应用。项目当前在 GitHub 上已收获超 1.1 万 Star,社区活跃,持续迭代。

如果你正在寻找一款免费、可本地部署的 AI 语音全流程工具,不妨立即访问 Voice-Pro GitHub 仓库,体验从视频到多语言配音的一站式工作流。

English

Voice-Pro: All-in-One AI Voice Processing Toolkit, Full Workflow from Recognition to Dubbing

In the field of voice processing, creators typically need to switch between multiple tools: downloading videos, separating vocals, transcribing text, translating subtitles, and then synthesizing dubbing. Voice-Pro integrates all these steps into a ready-to-use web application — built on Gradio, it supports YouTube video downloads, vocal separation, speech recognition, translation, subtitle generation, and multilingual dubbing, providing content creators, researchers, and cross-border teams with a complete AI voice processing pipeline.

Key Features: Five Core Capabilities Covering the Full Voice Processing Chain

  • Top-Tier Speech Recognition: Integrates three engines — Whisper, Faster-Whisper, and Whisper-Timestamped — supporting high-precision transcription with word-level timestamps, delivering excellent performance in Chinese, English, and multilingual recognition.
  • Zero-Shot Voice Cloning: Built-in F5-TTS, E2-TTS, and CosyVoice (including Fun-CosyVoice3 supporting 9 languages such as Korean), allowing you to clone any voice timbre without training samples.
  • Multilingual Text-to-Speech: Based on Edge-TTS and kokoro, with optional Azure TTS integration, generating natural and fluent multilingual dubbing with a single click.
  • YouTube Video Processing: Built-in downloader (yt-dlp) that can directly fetch videos for subtitle generation, translation, and re-dubbing to create multilingual content.
  • Professional Vocal Separation: Integrates Demucs and MDX-Net models, supporting accompaniment/vocal separation, suitable for karaoke, podcast editing, and mixing scenarios.

Highlights: Open Source, Local Running, Privacy and Security

Voice-Pro is released under the LGPL open-source license. All models and processing run entirely locally (Windows platform, with CUDA acceleration support), so user data never needs to be uploaded to the cloud — privacy and security remain fully under your control. Compared to commercial dubbing SaaS solutions, it delivers near ElevenLabs-level cloning and dubbing capabilities without monthly subscriptions — deploy once and use it long-term. For batch processing scenarios, its pipeline-style task design also significantly boosts efficiency.

Who It's For: Who Is Voice-Pro Best Suited For?

Content creators can use it to batch-generate multilingual subtitles and dubbing for rapid global expansion; podcasters and video editors can efficiently process material with vocal separation and transcription features; language learners can immerse themselves through video translation and bilingual subtitles; developers can build their own voice applications by integrating its modular architecture. The project has already garnered over 11,000 stars on GitHub, with an active community and continuous iteration.

If you're looking for a free, locally deployable AI voice processing tool covering the entire workflow, visit the Voice-Pro GitHub repository now and experience the one-stop workflow from video to multilingual dubbing.

滚动至顶部