qwenlm/qwen3-omni

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as...

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 1 hour ago
Added to GitGenius on September 17th, 2026
Created on September 21st, 2025
Open Issues & Pull Requests: 8 (+0)
GitHub issues: Enabled
Number of forks: 297
Total Stargazers: 4,019 (+0)
Total Subscribers: 26 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 39.4 hours
Mean response time: 11.3 days
90th percentile: 40.7 days
Tracked items: 32

Most active contributors

Sign in to see contributor activity.

How this project is maintained

Roughly one issue in three opened in the past year never receives a reply. 93% of issues opened in the past year have since been closed. Three people close 69% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 3
New in 7 days: 1
Closed in 7 days: 1
Avg open age: 35 days
Stale 30+ days: 0
Stale 90+ days: 0

Recent activity

Opened in 7 days: 1
Closed in 7 days: 1
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • inactive (30)

Most active issues this week

Detailed Description

Qwen3-Omni is a natively end-to-end omni-modal large language model that processes text, audio, images, and video while generating real-time streaming speech responses.

The model addresses the challenge of building unified systems that handle multiple input and output modalities without requiring separate specialized components. Rather than treating different modalities as separate problems, Qwen3-Omni integrates them into a single end-to-end architecture designed to understand diverse inputs and produce natural speech output in real time.

Developers should consider this tool if they need a single model capable of handling multimodal inputs across text, images, audio, and video with streaming speech generation. It suits applications requiring unified processing of mixed-media content without the complexity of chaining multiple specialized models. The project provides access through multiple channels including a web chat interface, Hugging Face and ModelScope repositories, local deployment options via Transformers and vLLM, and an API service through DashScope. Cookbooks are available to guide implementation for specific use cases.

Maintainers typically respond to new issues and pull requests within a few days. Work in the issue tracker is dominated by inactive items.