jamiepine/voicebox

Description: The open-source AI voice studio. Clone, dictate, create.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 32 minutes ago
Added to GitGenius on June 30th, 2026
Created on January 25th, 2026
Open Issues & Pull Requests: 576 (+0)
Number of forks: 5,308
Total Stargazers: 43,509 (+1)
Total Subscribers: 211 (+0)

Issue Activity (beta)

Open issues: 454
New in 7 days: 14
Closed in 7 days: 2
Avg open age: 71 days
Stale 30+ days: 348
Stale 90+ days: 185

Recent activity

Opened in 7 days: 12
Closed in 7 days: 2
Comments in 7 days: 3
Events in 7 days: 6

Top labels

  • bug (54)
  • enhancement (35)
  • invalid (4)
  • question (3)
  • good first issue (2)
  • documentation (1)
  • duplicate (1)

Most active issues this week

Repository Insights (GitGenius)

Median issue/PR response: 10.6 hours
Mean response time: 5.9 days
90th percentile: 16.1 days
Tracked items: 447

Most active contributors

Detailed Description

Voicebox is an open-source, local-first AI voice studio built with TypeScript that provides a complete voice input and output stack running entirely on users' machines. The project serves as a privacy-focused alternative to cloud-based services like ElevenLabs and WisprFlow, combining voice cloning, text-to-speech generation, and speech-to-text dictation into a single application. The repository is maintained primarily by jamiepine, who has logged 372 tracked events, with additional contributions from Vishaldadlani321 and sumire608.

The application supports seven distinct text-to-speech engines with different capabilities and language coverage. Qwen3-TTS and Qwen CustomVoice handle 10 languages with high-quality multilingual cloning and delivery instruction support. LuxTTS provides lightweight English synthesis at 48kHz output. Chatterbox Multilingual covers the broadest language range with 23 languages including Arabic, Danish, Finnish, Greek, Hebrew, Hindi, Malay, Norwegian, Polish, Swahili, Swedish, and Turkish. Chatterbox Turbo offers fast English synthesis with paralinguistic emotion and sound tags. HumeAI's TADA engine supports 10 languages with extended coherent audio generation up to 700 seconds. Kokoro provides 50 curated preset voices using a tiny 82-million parameter model optimized for CPU inference.

Voice cloning functionality includes zero-shot cloning from audio samples and access to over 50 curated preset voices. The application features post-processing audio effects including pitch shifting, reverb, delay, chorus, flanger, compression, gain adjustment, and high-pass and low-pass filtering. Users can create reusable effect presets and assign defaults per voice profile. Generation supports unlimited text length through automatic sentence-boundary splitting with crossfading, handling up to 50,000 characters with configurable chunk sizes between 100 and 5,000 characters.

The Stories Editor enables multi-track timeline composition for conversations, podcasts, and narratives with drag-and-drop functionality, inline audio trimming, and synchronized playback. Voice profiles support creation from audio files or direct in-app recording, with import and export capabilities for sharing and backup. The application includes a global dictation hotkey for system-wide voice input with push-to-talk and toggle modes, plus Whisper-based speech-to-text transcription.

Integration capabilities include a REST API and built-in MCP server for connecting voice I/O to external applications and AI agents. The application is built with Tauri using Rust for native performance rather than Electron, with platform-specific optimizations including MLX and Metal acceleration on macOS, CUDA on Windows, and AMD ROCm and Intel Arc support. Docker deployment is available alongside native installers for macOS, Windows, and Linux.

According to GitGenius tracking data, the repository has grown from 37,660 to 37,661 stargazers since July 4, 2026. Issue and pull request response latency shows a median of 10.4 hours and mean of 129.3 hours across 426 tracked items. Bug reports represent the most active issue category with 54 items, followed by enhancement requests with 35 items. The project maintains connections with related repositories including nousresearch/hermes-agent, anomalyco/opencode, and openclaw/openclaw through overlapping contributor networks.

voicebox
by
jamiepinejamiepine/voicebox

Repository Details

Fetching additional details & charts...