abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio...

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 24 minutes ago
Added to GitGenius on September 4th, 2026
Created on July 29th, 2024
Open Issues & Pull Requests: 60 (+0)
GitHub issues: Enabled
Number of forks: 1,844
Total Stargazers: 12,763 (+0)
Total Subscribers: 80 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 5.4 hours
Mean response time: 24.4 days
90th percentile: 58.1 days
Tracked items: 43

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 95% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 3% of issues opened in the past year have been closed. Three people close 93% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 37
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 387 days
Stale 30+ days: 34
Stale 90+ days: 31

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

Voice-Pro is a Gradio web application for speech recognition, translation, and AI voice synthesis.

The tool addresses the need for comprehensive audio processing and multilingual content creation by combining multiple specialized models in a single interface. It handles speech-to-text transcription using Whisper and Faster-Whisper, converts text to speech through Edge-TTS and kokoro, and performs zero-shot voice cloning via E2-TTS, F5-TTS, and CosyVoice. Additional capabilities include YouTube video download and processing, vocal isolation through Demucs, and multilingual translation. The application runs as a web interface accessible through a browser.

Creators and developers working with video content, podcasts, or audiobooks will find this tool useful for workflows involving transcription, dubbing, and subtitle generation. It suits projects requiring multilingual support and voice cloning without needing separate proprietary services. The tool is designed for Windows platforms and leverages CUDA for GPU acceleration where available.

The project maintains active development with regular updates and documentation in multiple languages. The codebase shows ongoing refinement of its audio processing pipeline and integration of multiple speech synthesis models. Community engagement is supported through documentation channels and external resources.