TMElyralab/MuseTalk

MuseTalk: Real-Time High Quality Lip Synchorization with Latent Space Inpainting

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 33 minutes ago
Added to GitGenius on September 10th, 2026
Created on March 26th, 2024
Open Issues & Pull Requests: 173 (+0)
GitHub issues: Enabled
Number of forks: 949
Total Stargazers: 6,549 (+0)
Total Subscribers: 68 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 43.4 hours
Mean response time: 18.1 days
90th percentile: 46.8 days
Tracked items: 210

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 100% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. Only 3% of issues opened in the past year have been closed. Three people close 58% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 133
New in 7 days: 0
Closed in 7 days: 0
Avg open age: 537 days
Stale 30+ days: 133
Stale 90+ days: 124

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

No label distribution available yet.

Most active issues this week

No issue events were indexed in the last 7 days.

Detailed Description

MuseTalk is a real-time audio-driven lip-syncing model that modifies facial video to match input audio with high fidelity.

The tool addresses the problem of creating convincing lip-synchronized video for virtual humans and dubbed content. It works by operating in the latent space of a frozen VAE encoder, taking audio input processed through a Whisper model and fusing audio embeddings with image features via cross-attention mechanisms borrowed from Stable Diffusion's UNet architecture. The approach enables real-time inference at 30 frames per second on standard GPU hardware while maintaining visual quality and synchronization accuracy.

MuseTalk suits projects requiring virtual human generation, video dubbing, or lip-sync correction where real-time performance matters. It handles multilingual audio including Chinese, English, and Japanese, and works with unseen faces at 256x256 resolution. The tool integrates with other virtual human pipelines, particularly MuseV for complete end-to-end solutions. Developers can train custom models using the provided training scripts and configurations, with checkpoints available trained on public and private datasets.

The project maintains active development with regular releases introducing significant improvements. Training code was open-sourced alongside version 1.5, which incorporated perceptual loss, GAN loss, and sync loss to enhance overall performance. A two-stage training strategy and spatio-temporal data sampling approach were implemented to balance visual quality against lip-sync accuracy. Pretrained model weights are available, and the team provides both inference and training code alongside a Gradio demo for experimentation.