openmoss/moss-tts

An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 4 minutes ago
Added to GitGenius on September 16th, 2026
Created on February 7th, 2026
Open Issues & Pull Requests: 19 (+0)
GitHub issues: Enabled
Number of forks: 373
Total Stargazers: 4,121 (+0)
Total Subscribers: 21 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 15.5 hours
Mean response time: 5.0 days
90th percentile: 11.5 days
Tracked items: 76

Most active contributors

Sign in to see contributor activity.

How this project is maintained

About 6% of issues opened in the past year have never received a reply. 81% of issues opened in the past year have been closed, leaving a working backlog. Three people close 82% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 15
New in 7 days: 1
Closed in 7 days: 0
Avg open age: 69 days
Stale 30+ days: 12
Stale 90+ days: 1

Recent activity

Opened in 7 days: 1
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • help wanted (1)
  • question (1)

Most active issues this week

Detailed Description

MOSS-TTS is a text-to-speech model family that generates long-form speech, dialogue, voice designs, sound effects, and supports real-time streaming synthesis.

The project addresses the need for flexible speech synthesis beyond simple sentence-level generation. It handles extended audio production including dialogue between multiple speakers, custom voice design, and sound effect generation. The approach leverages a model family architecture that supports both streaming and non-streaming inference modes, enabling real-time applications while maintaining quality for longer-form content generation.

The tool suits developers building voice applications that require more than basic TTS functionality, particularly those needing dialogue synthesis with multiple speakers or custom voice characteristics. Projects involving interactive voice systems, audio content creation, or applications demanding low-latency speech output would benefit from the streaming capabilities. The multilingual support extends its applicability across different language markets.

The project shows active development with regular commits and ongoing refinement of its model architecture. Work continues on expanding the capabilities of the model family to handle increasingly complex audio synthesis tasks. The codebase receives consistent updates addressing both core functionality and user-facing features.