Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-24 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-23 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-22 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-21 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-20 | 1 | 0 | 0 | 0 | 2 |
| 2026-08-19 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-18 | 0 | 0 | 2 | 2 | 2 |
| 2026-08-17 | 2 | 0 | 0 | 0 | 3 |
| 2026-08-16 | 1 | 0 | 0 | 0 | 0 |
| 2026-08-15 | 1 | 0 | 0 | 0 | 1 |
| 2026-08-14 | 0 | 1 | 0 | 0 | 0 |
| 2026-08-13 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-12 | 1 | 0 | 0 | 0 | 3 |
| 2026-08-11 | 1 | 1 | 0 | 0 | 0 |
Opened: 1
Closed: 0
Comments: 2
Events: 2
No label distribution is available yet.
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#2811 tests/test_flash_attn.py fails collection without einops Opened 4 days ago | nm-z | open | No labels | 0 | 0 | 4 days ago |
#2307 FA4 SM120 support? Opened 6 months ago | lucellent | open | No labels | 6 | 0 | 5 days ago |
#2581 Support head_dim=512 on SM89 (Ada) for Gemma 4 global attention layers Opened 3 months ago | ShuaiShao93 | open | No labels | 3 | 0 | 6 days ago |
#2799 Flash Attention install with Triton backend replaces preinstalled Triton no matter what. Opened 8 days ago | HimeLexie | open | No labels | 2 | 0 | 6 days ago |
#2801 [FA4] max_seqlen_q tensor causes an argument-annotation mismatch in varlen forward Opened 7 days ago | TorinLi | open | No labels | 0 | 0 | 7 days ago |
#2800 [ROCm][CK][gfx1201][Windows] FlashAttention backward is 4.4–7.4x slower than PyTorch SDPA on RX 9070 XT Opened 7 days ago | qweqweewqe7-create | open | No labels | 0 | 0 | 7 days ago |
#2797 FA4 backward does not return against cutlass-dsl 4.6.0 and 4.6.1 on B300 (sm_103) Opened 9 days ago | andylizf | open | No labels | 0 | 0 | 9 days ago |
#2438 pack_gqa fails on SM120 (SM80 CuTe DSL path): crd2idx coordinate resolution error Opened 5 months ago | blake-snc | closed - completed | No labels | 4 | 2 | 10 days ago |
#2793 Varlen version of `flash_attn_with_kvcache` Opened 12 days ago | ir2718 | open | No labels | 0 | 0 | 12 days ago |
#2769 publish.yml: create trigger has no tag filter, so branch pushes fail the workflow and new matrix entries never build Opened 17 days ago | joanfabregat | open | No labels | 1 | 0 | 13 days ago |
#2782 CuTe DSL: replace removed quack.activation.sub_packed_f32x2 API Opened 14 days ago | xihucuyu888 | closed - completed | No labels | 0 | 0 | 13 days ago |
#2789 ops/triton/linear.py kernel_fwd computes its bounds mask and never passes it to the store Opened 13 days ago | truong-v | open | No labels | 0 | 0 | 13 days ago |
#2783 raise error cuda12.6 pytorch 2.10, vllm0.18.0 when install flash-attn==2.8.1 Opened 14 days ago | cqray1990 | open | No labels | 0 | 0 | 14 days ago |
#2765 [CuTe, SM100] test_clc_fuzz fails after #2559 due to stale scheduler expectations Opened 17 days ago | JiaxuanBai | open | No labels | 0 | 0 | 17 days ago |
#2442 [Wheel] Pre-built flash_attn v2.8.3 for CUDA 13 + PyTorch 2.11 (Python 3.12) Opened 5 months ago | adithyaxx | open | No labels | 3 | 17 | 18 days ago |
#2764 [Cute, SM100] FA4 and FA4_Varlen fwd runtime difference reaches 27% Opened 18 days ago | complexfilter | open | No labels | 0 | 0 | 18 days ago |
#2760 [CuTe, SM100] Deadlock in varlen + block-sparse + SplitKV forward: packed `split_idx` is not unpacked on the block-sparse paths (regression from #2559) Opened 18 days ago | JiaxuanBai | open | No labels | 0 | 0 | 18 days ago |
#1981 Benchmarking FA on 5070Ti Opened 10 months ago | vin-prabhakar | open | No labels | 1 | 0 | 19 days ago |
#2699 [FA4] FA4 performs about the same as FA2 on B300. Opened 2 months ago | hewangsh | open | No labels | 6 | 0 | 20 days ago |
#2725 FlashAttention-2 on AMD Radeon RX 9060 XT / gfx1200 with ROCm 7.14 Opened 1 month ago | dreamdayapps | open | No labels | 1 | 0 | 22 days ago |
#1810 FA3 attention sinks blackwell sm120 Opened 1 year ago | fernandaspets | open | No labels | 7 | 0 | 24 days ago |
#2750 flash_attn_func crashes with cudaErrorInvalidValue (SMEM overflow) at head_dim=256, causal, bf16 on sm_86/sm_89 (SM80 tile heuristic not head-dim-aware) Opened 24 days ago | arbi-dev | open | No labels | 0 | 0 | 24 days ago |
#2425 [Wheel] Pre-built flash-attn 2.8.3 for CUDA 12 + PyTorch 2.11 (Python 3.10-3.13) Opened 5 months ago | lesj0610 | open | No labels | 7 | 25 | 25 days ago |
#2649 FlashAttentionForwardSm120: use_tma_O incorrectly True on SM_121 → AttributeError on tma_atom_O=None Opened 2 months ago | TyGu1 | closed - completed | No labels | 4 | 0 | 26 days ago |
#1632 Clarification on autotune using the triton backend for amd cards Opened 1 year ago | Kademo15 | open | No labels | 2 | 3 | 27 days ago |