jamiepine/voicebox

The open-source AI voice studio. Clone, dictate, create.

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 38 minutes ago
Added to GitGenius on June 30th, 2026
Created on January 25th, 2026
Open Issues & Pull Requests: 642 (+0)
Number of forks: 6,260
Total Stargazers: 50,493 (+3)
Total Subscribers: 231 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 20.9 hours
Mean response time: 9.5 days
90th percentile: 30.5 days
Tracked items: 415

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 55% of tracked open issues have had no activity in three months. Only 7% of issues opened in the past year have been closed. Three people close 75% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 489
New in 7 days: 17
Closed in 7 days: 2
Avg open age: 68 days
Stale 30+ days: 407
Stale 90+ days: 259

Recent activity

Opened in 7 days: 12
Closed in 7 days: 1
Comments in 7 days: 9
Events in 7 days: 13

Top labels

  • bug (54)
  • enhancement (35)
  • invalid (4)
  • question (3)
  • good first issue (2)
  • documentation (1)
  • duplicate (1)

Detailed Description

Voicebox is an open-source, local-first AI voice studio built with TypeScript that provides a complete voice input and output stack running entirely on users' machines. The project serves as a privacy-focused alternative to cloud-based services like ElevenLabs and WisprFlow, combining voice cloning, text-to-speech generation, and speech-to-text dictation into a single application.

The application supports seven distinct text-to-speech engines with different capabilities and language coverage. Qwen3-TTS and Qwen CustomVoice handle 10 languages with high-quality multilingual cloning and delivery instruction support. LuxTTS provides lightweight English synthesis at 48kHz output. Chatterbox Multilingual covers the broadest language range with 23 languages including Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, and Turkish. Chatterbox Turbo offers fast English synthesis with paralinguistic emotion and sound tags. HumeAI's TADA engine supports 10 languages with extended coherent audio generation up to 700 seconds. Kokoro provides 50 curated preset voices using a tiny 82-million parameter model optimized for CPU inference.

Voice cloning functionality includes zero-shot cloning from audio samples and access to over 50 curated preset voices. The application features post-processing audio effects including pitch shifting, reverb, delay, chorus, flanger, compression, gain adjustment, and high-pass and low-pass filtering. Users can create reusable effect presets and assign defaults per voice profile. Generation supports unlimited text length through automatic sentence-boundary splitting with crossfading, handling up to 50,000 characters with configurable chunk sizes between 100 and 5,000 characters.

The Stories Editor enables multi-track timeline composition for conversations, podcasts, and narratives with drag-and-drop functionality, inline audio trimming, and synchronized playback. Voice profiles support creation from audio files or direct in-app recording, with import and export capabilities for sharing and backup. The application includes a global dictation hotkey for system-wide voice input with push-to-talk and toggle modes, plus Whisper-based speech-to-text transcription.

Integration capabilities include a REST API and built-in MCP server for connecting voice I/O to external applications and AI agents. The application is built with Tauri using Rust for native performance rather than Electron, with platform-specific optimizations including MLX and Metal acceleration on macOS, CUDA on Windows, and AMD ROCm and Intel Arc support. Docker deployment is available alongside native installers for macOS, Windows, and Linux.