paddlepaddle/paddlespeech

Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker...

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 5 minutes ago
Added to GitGenius on September 4th, 2026
Created on November 14th, 2017
Open Issues & Pull Requests: 276 (+0)
GitHub issues: Enabled
Number of forks: 1,959
Total Stargazers: 12,680 (+0)
Total Subscribers: 187 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 25.9 hours
Mean response time: 10.5 days
90th percentile: 9.0 days
Tracked items: 421

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 93% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Work labelled "S2T" is answered fastest, typically in under an hour, while "Question" waits about 33 hours. Only 5% of issues opened in the past year have been closed. Three people close 60% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 57
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 748 days
Stale 30+ days: 57
Stale 90+ days: 55

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • Question (335)
  • Stale (316)
  • T2S (36)
  • S2T (35)
  • Bug (32)
  • feature request (10)
  • Installation (3)
  • Report (2)

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

PaddleSpeech is a speech and audio toolkit built on the PaddlePaddle platform that provides end-to-end solutions for automatic speech recognition, text-to-speech synthesis, speaker verification, speech translation, and keyword spotting.

The toolkit addresses the need for accessible, production-ready speech processing by bundling state-of-the-art models with unified interfaces. It supports streaming automatic speech recognition with punctuation restoration, streaming text-to-speech synthesis with integrated text frontend processing, self-supervised learning models, speaker verification systems, end-to-end speech translation, and keyword spotting. The architecture emphasizes ease of use through pre-trained models and standardized APIs across different speech tasks.

Developers should choose this toolkit if they need multiple speech capabilities in a single framework rather than integrating separate specialized libraries. It suits projects requiring streaming speech recognition and synthesis, multilingual speech translation, or speaker identification. The toolkit runs on Linux, Windows, and macOS with Python support. The README does not make explicit comparisons to alternative speech toolkits, so no comparative positioning can be stated.

The project maintains active development with regular commits and ongoing issue engagement. The codebase shows consistent contribution activity from multiple developers. The toolkit has been recognized with a major award for its demonstration of practical speech processing capabilities.