Voice-Pro is a Gradio web application for speech recognition, translation, and AI voice synthesis.
The tool addresses the need for comprehensive audio processing and multilingual content creation by combining multiple specialized models in a single interface. It handles speech-to-text transcription using Whisper and Faster-Whisper, converts text to speech through Edge-TTS and kokoro, and performs zero-shot voice cloning via E2-TTS, F5-TTS, and CosyVoice. Additional capabilities include YouTube video download and processing, vocal isolation through Demucs, and multilingual translation. The application runs as a web interface accessible through a browser.
Creators and developers working with video content, podcasts, or audiobooks will find this tool useful for workflows involving transcription, dubbing, and subtitle generation. It suits projects requiring multilingual support and voice cloning without needing separate proprietary services. The tool is designed for Windows platforms and leverages CUDA for GPU acceleration where available.
The project maintains active development with regular updates and documentation in multiple languages. The codebase shows ongoing refinement of its audio processing pipeline and integration of multiple speech synthesis models. Community engagement is supported through documentation channels and external resources.