openbmb/voxcpm

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 8 minutes ago
Added to GitGenius on August 31st, 2026
Created on September 16th, 2025
Open Issues & Pull Requests: 112 (+0)
Number of forks: 4,165
Total Stargazers: 36,379 (-1)
Total Subscribers: 158 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 13.9 hours
Mean response time: 5.5 days
90th percentile: 8.4 days
Tracked items: 305

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 61% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 12% of issues opened in the past year have been closed. Three people close 75% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 93
New in 7 days: 0
Closed in 7 days: 2
Avg open age: 134 days
Stale 30+ days: 78
Stale 90+ days: 56

Recent activity

Opened in 7 days: 0
Closed in 7 days: 2
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Detailed Description

VoxCPM is a tokenizer-free text-to-speech system that generates multilingual speech synthesis with voice cloning and creative voice design capabilities.

The tool addresses the limitations of traditional TTS systems that rely on discrete tokenization by using an end-to-end diffusion autoregressive architecture to directly generate continuous speech representations. This approach eliminates the information loss and artifacts that can occur when speech is compressed into discrete tokens, enabling more natural and expressive synthesis. The system operates without requiring a separate tokenizer stage, streamlining the generation pipeline from text to high-quality audio output.

VoxCPM2 is suited for developers building multilingual speech applications, particularly those needing studio-quality audio at 48kHz across 30 languages. The tool supports controllable voice cloning, allowing users to synthesize speech in specific voices with fine-grained control, and voice design features for creative audio generation. Projects requiring natural-sounding, expressive speech synthesis with support for diverse languages and custom voice characteristics would benefit from this approach. The system is built on a 2 billion parameter model trained on extensive multilingual speech data, making it capable of handling complex linguistic and acoustic requirements.

The project maintains active development with regular updates to its codebase and documentation. The tool provides multiple access points including a live demonstration interface, comprehensive documentation, and model availability on standard machine learning platforms. Community engagement is facilitated through multiple channels for discussion and support.