ysharma3501/luxtts

A high-quality rapid TTS voice cloning model that reaches speeds of 150x realtime.

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 27 minutes ago
Added to GitGenius on September 12th, 2026
Created on January 23rd, 2026
Open Issues & Pull Requests: 32 (+0)
GitHub issues: Enabled
Number of forks: 682
Total Stargazers: 5,378 (+0)
Total Subscribers: 39 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 29.9 hours
Mean response time: 11.9 days
90th percentile: 32.1 days
Tracked items: 29

Most active contributors

Sign in to see contributor activity.

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 10% of issues opened in the past year have been closed. Three people close 67% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 25
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 175 days
Stale 30+ days: 24
Stale 90+ days: 19

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

Detailed Description

LuxTTS is a text-to-speech model designed for high-quality voice cloning that achieves synthesis speeds exceeding 150x realtime.

The tool addresses the challenge of performing realistic voice cloning efficiently by using a lightweight architecture based on ZipVoice principles but distilled to four inference steps with an improved sampling technique. It generates clear 48kHz speech output, a higher quality than most TTS models limited to 24kHz, while maintaining a memory footprint under 1GB of VRAM. The approach enables the model to run on consumer GPUs and even faster than realtime on CPUs.

Developers should choose this tool if they need voice cloning capabilities with minimal computational overhead and want to run inference locally without cloud dependencies. It suits projects requiring high-quality speech synthesis at scale or on resource-constrained hardware. The README distinguishes it from ZipVoice by highlighting the distillation to fewer steps, improved sampling, and the custom 48kHz vocoder. The tool requires a minimum 3-second audio file for voice cloning and offers tuning parameters like return_smooth for addressing metallic artifacts and t_shift for balancing pronunciation accuracy against output quality.

The project maintains an active ecosystem with community-contributed implementations including Gradio interfaces, ComfyUI nodes, and ONNX variants. Development roadmap items indicate planned releases for version 1.5 and float16 inference optimization, suggesting ongoing work to improve performance further. The codebase is available with pre-trained models hosted on Hugging Face alongside interactive demo spaces and Colab notebooks for immediate experimentation.