microsoft/VibeVoice

Open-Source Frontier Voice AI

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 20 minutes ago
Added to GitGenius on December 6th, 2025
Created on August 25th, 2025
Open Issues & Pull Requests: 183 (+0)
Number of forks: 5,996
Total Stargazers: 53,166 (+2)
Total Subscribers: 259 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 10.9 hours
Mean response time: 5.3 days
90th percentile: 14.5 days
Tracked items: 209

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 86% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 10% of issues opened in the past year have been closed. Three people close 71% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 128
New in 7 days: 1
Closed in 7 days: 0
Avg open age: 83 days
Stale 30+ days: 125
Stale 90+ days: 110

Recent activity

Opened in 7 days: 1
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • bug (1)
  • compatibility (1)

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

VibeVoice is an open-source voice AI framework that combines speech synthesis and speech recognition capabilities. The project provides tools for text-to-speech generation with emotional control and automatic speech recognition that can process long-form audio in a single pass.

The framework addresses the need for unified voice AI by offering both synthesis and recognition in one system. VibeVoice-TTS generates speech with controllable emotional characteristics, while VibeVoice-ASR handles speech-to-text conversion for extended audio, producing structured transcriptions that include speaker identification, timestamps, and content. The ASR component supports over fifty languages and can process up to sixty minutes of audio without segmentation, with optional user-customized context integration.

Developers should consider this tool if they need multilingual speech recognition or emotionally expressive text-to-speech in a single framework. The project suits applications requiring long-form audio processing without manual segmentation, or systems where emotional nuance in synthesized speech matters. The tool integrates with standard machine learning infrastructure, including support for vLLM inference acceleration and compatibility with the Hugging Face Transformers library.

The project maintains active engagement with its user base, with nearly all open issues originating from external adopters rather than the core team, indicating substantial real-world adoption. Maintainers respond to new issues and pull requests within a day. Work in the issue tracker centers on compatibility and bug resolution, reflecting a focus on stability and integration across different deployment environments.