0xShug0/audio.cpp

An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with...

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 37 minutes ago
Added to GitGenius on September 24th, 2026
Created on June 23rd, 2026
Open Issues & Pull Requests: 19 (+0)
GitHub issues: Enabled
Number of forks: 393
Total Stargazers: 3,399 (+0)
Total Subscribers: 32 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 1.1 hours
Mean response time: 9.7 hours
90th percentile: 9.7 hours
Tracked items: 261

How this project is maintained

Practically every issue opened in the past year has drawn a reply. 97% of issues opened in the past year have since been closed. Three people close 92% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 9
New in 7 days: 21
Closed in 7 days: 19
Avg open age: 23 days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 17
Closed in 7 days: 16
Comments in 7 days: 57
Events in 7 days: 132

Top labels

  • feature (27)
  • new model (26)
  • webui (8)
  • not reproducible (7)
  • showcase (5)
  • announcement (4)
  • bug (4)
  • good first issue (4)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

audio.cpp is a C++ inference engine for audio models that eliminates the need for Python dependencies and conda environments by providing a unified native runtime for speech and audio AI tasks.

The tool solves the friction of managing multiple Python environments and dependency conflicts when working with audio models. It builds on ggml to deliver a single C++ framework that handles text-to-speech, speech-to-text, voice conversion, voice cloning, music generation, diarization, voice activity detection, and source separation. The engine runs on Windows, Linux, and macOS with support for NVIDIA, AMD, and Apple Silicon GPUs as well as CPU-only inference.

Developers should choose this tool if they need to deploy audio models in production environments where Python overhead and dependency management are problematic, or if they want to experiment with multiple audio models without environment conflicts. It suits edge deployment, real-time applications, and scenarios where latency and resource efficiency matter. The project demonstrates significant performance gains over Python reference implementations, with some TTS paths running up to eight times faster and end-to-end latency reduced by forty-five to eighty-five percent on CUDA. Quantized GGUF models can run substantially faster while reducing peak memory usage. The tool includes a WebUI with an Arena tab for comparing models side by side.

The project shows active development with recent releases adding support for new model families and optimization work across multiple inference paths. Performance improvements are continuously measured and documented, with detailed reports comparing quantized and full-precision variants. The codebase maintains support across diverse hardware platforms and regularly integrates new audio model architectures into the framework.