NVlabs/Sana

Description: SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer

View on GitHub ↗Jump to charts ↓

Summary Information

Updated 59 minutes ago
Added to GitGenius on March 8th, 2026
Created on October 11th, 2024
Open Issues & Pull Requests: 135 (+0)
Number of forks: 685
Total Stargazers: 8,537 (+0)
Total Subscribers: 98 (+0)

Issue Activity (beta)

Open issues: 126
New in 7 days: 1
Closed in 7 days: 0
Avg open age: 221 days
Stale 30+ days: 121
Stale 90+ days: 98

Recent activity

Opened in 7 days: 0
Closed in 7 days: 0
Comments in 7 days: 0
Events in 7 days: 0

Top labels

  • Answered (115)
  • fixed (24)
  • bug (12)
  • documentation (10)
  • working (9)
  • Announcement (6)
  • enhancement (1)
  • pachage version bug (1)

Most active issues this week

No issue events were indexed in the last 7 days.

Repository Insights (GitGenius)

Median issue/PR response: 11.2 hours
Mean response time: 4.9 days
90th percentile: 19.0 days
Tracked items: 247

Most active contributors

Detailed Description

SANA is an efficiency-oriented codebase developed by NVIDIA Labs for high-resolution image and video generation, providing complete training and inference pipelines. The repository implements a family of models including SANA, SANA-1.5, SANA-Sprint, SANA-Video, SANA-WM, SANA-Streaming, and Sol-RL, all built around the core concept of linear diffusion transformers for efficient synthesis. The project has achieved significant academic recognition, with papers accepted as oral presentations at ICLR 2025 and ICLR 2026, as well as highlights at ICCV 2025 and acceptance at ICML 2025.

The repository centers on linear transformer architectures as an alternative to standard diffusion models, enabling efficient generation at high resolutions. SANA-Video, released in October 2025, supports both text-to-video and text-image-to-video generation with a 5-second linear DiT video model, and has been integrated into the Hugging Face diffusers library. SANA-WM represents a 2.6 billion parameter controllable world model supporting 720p, 1-minute video generation with 6-degree-of-freedom camera control, positioning it as a baseline for world modeling and embodied AI applications. SANA-Streaming, released in June 2026, is a 2B model for real-time streaming video editing at 720p resolution for up to one-minute videos, marking a pioneering approach to streaming editing workflows.

The codebase includes Sol-RL, which provides reinforcement learning infrastructure with NVFP4 rollout capabilities and BF16 training support, offering complete training recipes for SANA, FLUX.1, and SD3.5-L models bundled with post-training datasets. The project has established partnerships with major frameworks, including integration with SGLang for high-performance serving with OpenAI-compatible APIs and collaboration with Cosmos-RL to provide complete RL infrastructure for post-training SANA models using state-of-the-art algorithms like Diffusion-NFT and Flow-GRPO.

According to GitGenius activity tracking across 247 issues and pull requests, the repository maintains a median response latency of 11.2 hours with a mean of 118.3 hours, indicating active community engagement. The most frequently applied issue labels are Answered (115 instances), fixed (24), and bug (12), reflecting a well-maintained codebase. Primary contributors include lawrence-cj with 611 tracked events, yujincheng08 with 53 events, and nitinmukesh with 51 events. The repository shares overlapping contributors with major projects including Hugging Face diffusers, PyTorch, and Hugging Face transformers, indicating deep integration within the broader deep learning ecosystem.

The codebase is classified across multiple technical domains including neural architecture search, deep learning, neural networks, architecture optimization, and scalable high-performance systems. The project emphasizes system-algorithm design and model discovery, with topics spanning diffusion models, linear transformers, text-to-image generation, text-to-video synthesis, streaming video, reinforcement learning, and world models. Multiple deployment options are available, including demos on dedicated hardware configurations, Hugging Face spaces, Replicate API services, and community integrations through ComfyUI. The project maintains active community engagement through a Discord server and comprehensive documentation accessible through its official website.

Sana
by
NVlabsNVlabs/Sana

Repository Details

Fetching additional details & charts...