Repository Issue Activity (beta)

flashinfer-ai/flashinfer

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
326
New in 7 Days
36
Closed in 7 Days
39
Average Open Age
98 days
Stale 30+ Days
152
Stale 90+ Days
66
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-121151420
2026-09-11101132423
2026-09-1062422325
2026-09-09711000
2026-09-08312000
2026-09-0782000
2026-09-0610000
2026-09-0500000
2026-09-04121000
2026-09-03112000
2026-09-0268000
2026-09-0132000
2026-08-3134000
2026-08-3000000
This Week

Opened: 36

Closed: 39

Comments: 12

Events: 60

Top Labels
needs-triage (428)
op: attention (182)
bug (176)
op: moe (139)
priority: must have (P0) (126)
feature request (115)
question (74)
op: gemm (59)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#5107 Split-KV decode rounds partial outputs to the output dtype, in a workspace already sized for fp32

Opened 2 days ago
Wint3rNight
open
bug
needs-triage
1010 hours ago

#2861 mm_fp4 trtllm backend leaks padding scales into real rows (use_8x4_sf_layout=True)

Opened 6 months ago
elvircrn
open
priority: must have (P0)
op: gemm
1011 hours ago

#3864 [Bug] trtllm-gen spec-as-decode (q_len_per_req=8) with BF16 query + FP8 KV cache silently mis-computes sliding-window attention on SM100

Opened 2 months ago
elad-inferize
open
needs-triage
3111 hours ago

#4254 [CAKE] Long-term CAKE-generated Kernel Progress Tracker

Opened 1 month ago
yyihuang
open
needs-triage
op: misc
op: linear attention
52611 hours ago

#5053 trtllm-gen NVFP4 MoE: the SiTU (siTuGlu) GEMM1 epilogue writes no output block scale factors for some configurations

Opened 4 days ago
kzjeef
open
needs-triage
1011 hours ago

#2511 [Bug] Cuda Graph Issues with TRTLLM-GEN Backend with BatchDecodeWithPagedKVCacheWrapper

Opened 7 months ago
NihalPotdar
closed - not_planned
bug
op: attention
4011 hours ago

#5058 top_k_page_table_transform(dsa_graph_safe=True) fails with cudaErrorNotSupported on SM120 (RTX 6000D), graph_safe=False works even under CUDA graph capture

Opened 3 days ago
qqtang-code
open
needs-triage
2016 hours ago

#3617 [feat] support fp8 per-token-head kv cache with inline scale

Opened 3 months ago
ir1ka
open
needs-triage
1016 hours ago

#5162 [Bug] Nightly Release: BF16 rank-major session CPU tests assume a source checkout

Opened 16 hours ago
cindyzxq
open
needs-triage
0016 hours ago

#5157 H100 CI broken on main since #5131: cudnn-frontend 1.28 vs backend 9.24 rejects non-ragged Stats layout (134 nodes, blocks all PRs)

Opened 20 hours ago
aleozlx
open
needs-triage
1020 hours ago

#5156 [Feature]: CuTe DSL W4A8 MXFP4 SiTU MoE for Kimi K3

Opened 20 hours ago
henrylhtsang
open
feature request
needs-triage
0020 hours ago

#5095 [Bug] SM120 sparse-MLA prefill (DSv4 dual-cache) returns corrupted output for >64 query tokens; masking lanes beyond topk_length to -1 fixes it

Opened 2 days ago
caiovicentino
closed - completed
needs-triage
1020 hours ago

#4973 [Bug] DSV4 Vision: illegal memory access on SM120 (2x RTX PRO 6000, TP=2) with long text prompts

Opened 8 days ago
JaviAFKzX
open
op: attention
2021 hours ago

#3620 [Bug] BatchPrefillWithPagedKVCache fails on SM75 (Turing / Tesla T4) with CUDA "invalid argument"

Opened 3 months ago
mikekg
closed - completed
needs-triage
priority: should have (P1)
1221 hours ago

#4936 [KDA] Unify the public `recurrent_kda` API, and settle the unreleased KDA surface before FI v0.7

Opened 9 days ago
kahyunnam
open
needs-triage
op: linear attention
3021 hours ago

#3800 [Bug] cudnn_batch_prefill_with_kv_cache: one-token Q batch stride silently corrupts every batch after the first when batch_offsets_q is not passed

Opened 2 months ago
waynehacking8
closed - completed
needs-triage
1021 hours ago

#3511 Trtllm-gen kernels not found: `headDim=512`, `tileSizeQ=128`

Opened 3 months ago
akelch11
closed - completed
needs-triage
1021 hours ago

#3971 SM103 serving hang caused by TRTLLM_GEN_BMM artifact regeneration in #3708 (batched_gemm-dd6d23e-721ae60, 0.6.14): FP4 batched-GEMM clusters stuck in mbarrier phase wait

Opened 2 months ago
YAMY1234
closed - completed
needs-triage
2021 hours ago

#4930 [Feature]: On SM107, the MXFP8 Quantization Only does not support FP32

Opened 9 days ago
Vinnie6167
closed - completed
feature request
needs-triage
priority: should have (P1)
0022 hours ago

#4999 [Bug] test_sparse_mla_sm120_cpb_model assert 0.0042064361572265625 <= (1.25 * 0.0028502678871154784)

Opened 6 days ago
nvamyt
closed - completed
needs-triage
4022 hours ago

#4957 [Bug] test_multistream_overlap.py CUDA error: unspecified launch failure Search for `cudaErrorLaunchFailure'

Opened 8 days ago
nvamyt
open
arch: sm107
1022 hours ago

#5151 [Feature]: MNNVL AllReduce does not support TriggerCompletionAtEnd

Opened 22 hours ago
benchislett
open
feature request
needs-triage
0022 hours ago

#5141 [Feature]: Expose NVFP4 4over6 configuration through the public API

Opened 1 day ago
pst2154
open
feature request
0024 hours ago

#4367 [Bug][v0.6.17] tests.trace.test_template_registry::test_registration_module_inventory_is_complete AssertionError: assert {'flashinfer....cat_ops', ...} == {'flashinfer....cat_ops', ...}

Opened 1 month ago
nvamyt
closed - completed
bug
needs-triage
ci: health
101 day ago

#5145 [Bug] JIT sampler built with CUDA 13 is reused by CUDA 12.9 Ray workers and fails at runtime

Opened 1 day ago
ZFXzzz
open
needs-triage
201 day ago

Rows per page:

1–25 of 1,238