collabora/whisperlive

A nearly-live implementation of OpenAI's Whisper.

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 2 hours ago
Added to GitGenius on September 16th, 2026
Created on May 4th, 2023
Open Issues & Pull Requests: 38 (+0)
GitHub issues: Enabled
Number of forks: 594
Total Stargazers: 4,290 (+0)
Total Subscribers: 42 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 3.6 days
Mean response time: 76.7 days
90th percentile: 406.6 days
Tracked items: 171

Most active contributors

Sign in to see contributor activity.

How this project is maintained

84% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Three people close 92% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 25
New in 7 days: 3
Closed in 7 days: 0
Avg open age: 592 days
Stale 30+ days: 21
Stale 90+ days: 20

Recent activity

Opened in 7 days: 3
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • Feature Request (17)
  • bug (7)
  • good first issue (2)
  • macos (2)
  • question (2)
  • windows (2)
  • help wanted (1)

Most active issues this week

Sign in to see which issues are moving.

Detailed Description

Whisperlive is a speech-to-text tool that implements OpenAI's Whisper model with near-real-time performance.

The tool addresses the latency problem in running Whisper for live transcription scenarios. Rather than processing audio in batches after collection, whisperlive streams audio input and produces transcription output with minimal delay. It achieves this by integrating multiple inference backends including TensorRT, OpenVINO, and ROCm, allowing users to optimize for their available hardware. The implementation supports both transcription and translation tasks, converting spoken audio directly to text in the source language or in English.

Whisperlive suits developers building live captioning systems, voice dictation applications, or real-time transcription features where latency matters. It works with OBS for streaming scenarios and supports text-to-speech output alongside transcription. The project is particularly valuable for those with GPU acceleration available, as the backend options target NVIDIA, Intel, and AMD hardware. Developers without specialized hardware can still use the tool but should expect different performance characteristics.

The project shows consistent maintenance with regular updates addressing both bug fixes and feature additions. Development activity includes ongoing refinement of inference backend integration and expansion of supported hardware platforms. The codebase receives attention to performance optimization, reflecting the core goal of achieving near-live transcription speeds. Community contributions are integrated regularly, and the project maintains responsiveness to reported issues and feature requests.