Dao-AILab/flash-attention

Fast and memory-efficient exact attention

View on GitHub ↗Jump to charts ↓Open shareable report

Summary Information

Updated 55 minutes ago
Added to GitGenius on February 25th, 2026
Created on May 19th, 2022
Open Issues & Pull Requests: 1,253 (+0)
Number of forks: 2,984
Total Stargazers: 24,708 (+0)
Total Subscribers: 149 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 14.9 hours
Mean response time: 60.7 days
90th percentile: 229.1 days
Tracked items: 1,039

How this project is maintained

Around half of the issues opened in the past year never receive a reply. 99% of open issues come from outside the core team, so the backlog reflects real-world use rather than internal planning. 78% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 3% of issues opened in the past year have been closed.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 858
New in 7 days: 4
Closed in 7 days: 1
Avg open age: 455 days
Stale 30+ days: 829
Stale 90+ days: 795

Recent activity

Opened in 7 days: 4
Closed in 7 days: 1
Comments in 7 days: 1
Events in 7 days: 2

Top labels

No label distribution available yet.

Detailed Description

Flash-Attention is an official implementation repository for FlashAttention and FlashAttention-2, algorithms designed to provide fast and memory-efficient exact attention computation with IO-awareness. The repository is written primarily in Python and addresses a critical bottleneck in transformer-based deep learning models by optimizing how attention mechanisms execute on GPUs.

The core contribution centers on scaled dot product attention computation, implementing the operation softmax(Q @ K^T * softmax_scale) @ V with significant performance improvements over standard PyTorch implementations. The repository provides multiple versions of the algorithm optimized for different hardware generations. FlashAttention-2 represents a complete rewrite that achieves approximately 2x speedup over the original version through improved parallelism and work partitioning. FlashAttention-3 is optimized specifically for Hopper GPUs such as the H100, currently available as a beta release supporting FP16, BF16 forward and backward passes, and FP8 forward passes on CUDA 12.3 and above. FlashAttention-4, written in CuTeDSL, extends support to both Hopper and Blackwell GPUs including the B200.

The repository maintains broad hardware compatibility across multiple platforms. For NVIDIA CUDA, it supports Ampere, Ada, and Hopper GPUs with datatypes fp16 and bf16, accommodating head dimensions up to 256. AMD ROCm support includes two backends: the Composable Kernel backend supporting MI200x, MI250x, MI300x, MI355x, and RDNA 3/4 GPUs, and a Triton backend supporting CDNA and RDNA GPUs with additional features like paged attention, rotary embeddings, and ALiBi. The codebase also references a separate flash-attention-turing repository for Turing GPU support.

Installation requires CUDA or ROCm toolkit, PyTorch 2.2 and above, and several Python packages including packaging, psutil, and ninja. The repository emphasizes that ninja installation is critical for compilation performance, reducing build time from approximately 2 hours to 3-5 minutes on 64-core machines. Windows support exists but requires additional testing.

The changelog documents substantial feature additions across versions. Version 2.1 modified causal masking behavior for unequal sequence lengths. Version 2.2 optimized inference scenarios with small query sequence lengths through KV cache handling. Version 2.3 introduced sliding window attention used in Mistral 7B. Version 2.4 added ALiBi support and deterministic backward passes. Version 2.5 implemented paged KV cache support via PagedAttention. Version 2.6 added softcapping for attention, used in Gemma-2 and Grok models. Version 2.7 achieved compatibility with torch compile.

The codebase is classified across multiple domains including attention mechanisms, transformers, GPU optimization, memory efficiency, performance acceleration, and large language models, reflecting its central role in modern AI training infrastructure.