Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-12 | 1 | 1 | 5 | 14 | 20 |
| 2026-09-11 | 10 | 11 | 3 | 24 | 23 |
| 2026-09-10 | 6 | 2 | 4 | 22 | 325 |
| 2026-09-09 | 7 | 11 | 0 | 0 | 0 |
| 2026-09-08 | 3 | 12 | 0 | 0 | 0 |
| 2026-09-07 | 8 | 2 | 0 | 0 | 0 |
| 2026-09-06 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-05 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-04 | 12 | 1 | 0 | 0 | 0 |
| 2026-09-03 | 11 | 2 | 0 | 0 | 0 |
| 2026-09-02 | 6 | 8 | 0 | 0 | 0 |
| 2026-09-01 | 3 | 2 | 0 | 0 | 0 |
| 2026-08-31 | 3 | 4 | 0 | 0 | 0 |
| 2026-08-30 | 0 | 0 | 0 | 0 | 0 |
Opened: 36
Closed: 39
Comments: 12
Events: 60
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#5107 Split-KV decode rounds partial outputs to the output dtype, in a workspace already sized for fp32 Opened 2 days ago | Wint3rNight | open | bug needs-triage | 1 | 0 | 10 hours ago |
#2861 mm_fp4 trtllm backend leaks padding scales into real rows (use_8x4_sf_layout=True) Opened 6 months ago | elvircrn | open | priority: must have (P0) op: gemm | 1 | 0 | 11 hours ago |
#3864 [Bug] trtllm-gen spec-as-decode (q_len_per_req=8) with BF16 query + FP8 KV cache silently mis-computes sliding-window attention on SM100 Opened 2 months ago | elad-inferize | open | needs-triage | 3 | 1 | 11 hours ago |
#4254 [CAKE] Long-term CAKE-generated Kernel Progress Tracker Opened 1 month ago | yyihuang | open | needs-triage op: misc op: linear attention | 5 | 26 | 11 hours ago |
#5053 trtllm-gen NVFP4 MoE: the SiTU (siTuGlu) GEMM1 epilogue writes no output block scale factors for some configurations Opened 4 days ago | kzjeef | open | needs-triage | 1 | 0 | 11 hours ago |
#2511 [Bug] Cuda Graph Issues with TRTLLM-GEN Backend with BatchDecodeWithPagedKVCacheWrapper Opened 7 months ago | NihalPotdar | closed - not_planned | bug op: attention | 4 | 0 | 11 hours ago |
#5058 top_k_page_table_transform(dsa_graph_safe=True) fails with cudaErrorNotSupported on SM120 (RTX 6000D), graph_safe=False works even under CUDA graph capture Opened 3 days ago | qqtang-code | open | needs-triage | 2 | 0 | 16 hours ago |
#3617 [feat] support fp8 per-token-head kv cache with inline scale Opened 3 months ago | ir1ka | open | needs-triage | 1 | 0 | 16 hours ago |
#5162 [Bug] Nightly Release: BF16 rank-major session CPU tests assume a source checkout Opened 16 hours ago | cindyzxq | open | needs-triage | 0 | 0 | 16 hours ago |
#5157 H100 CI broken on main since #5131: cudnn-frontend 1.28 vs backend 9.24 rejects non-ragged Stats layout (134 nodes, blocks all PRs) Opened 20 hours ago | aleozlx | open | needs-triage | 1 | 0 | 20 hours ago |
#5156 [Feature]: CuTe DSL W4A8 MXFP4 SiTU MoE for Kimi K3 Opened 20 hours ago | henrylhtsang | open | feature request needs-triage | 0 | 0 | 20 hours ago |
#5095 [Bug] SM120 sparse-MLA prefill (DSv4 dual-cache) returns corrupted output for >64 query tokens; masking lanes beyond topk_length to -1 fixes it Opened 2 days ago | caiovicentino | closed - completed | needs-triage | 1 | 0 | 20 hours ago |
#4973 [Bug] DSV4 Vision: illegal memory access on SM120 (2x RTX PRO 6000, TP=2) with long text prompts Opened 8 days ago | JaviAFKzX | open | op: attention | 2 | 0 | 21 hours ago |
#3620 [Bug] BatchPrefillWithPagedKVCache fails on SM75 (Turing / Tesla T4) with CUDA "invalid argument" Opened 3 months ago | mikekg | closed - completed | needs-triage priority: should have (P1) | 1 | 2 | 21 hours ago |
#4936 [KDA] Unify the public `recurrent_kda` API, and settle the unreleased KDA surface before FI v0.7 Opened 9 days ago | kahyunnam | open | needs-triage op: linear attention | 3 | 0 | 21 hours ago |
#3800 [Bug] cudnn_batch_prefill_with_kv_cache: one-token Q batch stride silently corrupts every batch after the first when batch_offsets_q is not passed Opened 2 months ago | waynehacking8 | closed - completed | needs-triage | 1 | 0 | 21 hours ago |
#3511 Trtllm-gen kernels not found: `headDim=512`, `tileSizeQ=128` Opened 3 months ago | akelch11 | closed - completed | needs-triage | 1 | 0 | 21 hours ago |
#3971 SM103 serving hang caused by TRTLLM_GEN_BMM artifact regeneration in #3708 (batched_gemm-dd6d23e-721ae60, 0.6.14): FP4 batched-GEMM clusters stuck in mbarrier phase wait Opened 2 months ago | YAMY1234 | closed - completed | needs-triage | 2 | 0 | 21 hours ago |
#4930 [Feature]: On SM107, the MXFP8 Quantization Only does not support FP32 Opened 9 days ago | Vinnie6167 | closed - completed | feature request needs-triage priority: should have (P1) | 0 | 0 | 22 hours ago |
#4999 [Bug] test_sparse_mla_sm120_cpb_model assert 0.0042064361572265625 <= (1.25 * 0.0028502678871154784) Opened 6 days ago | nvamyt | closed - completed | needs-triage | 4 | 0 | 22 hours ago |
#4957 [Bug] test_multistream_overlap.py CUDA error: unspecified launch failure Search for `cudaErrorLaunchFailure' Opened 8 days ago | nvamyt | open | arch: sm107 | 1 | 0 | 22 hours ago |
#5151 [Feature]: MNNVL AllReduce does not support TriggerCompletionAtEnd Opened 22 hours ago | benchislett | open | feature request needs-triage | 0 | 0 | 22 hours ago |
#5141 [Feature]: Expose NVFP4 4over6 configuration through the public API Opened 1 day ago | pst2154 | open | feature request | 0 | 0 | 24 hours ago |
#4367 [Bug][v0.6.17] tests.trace.test_template_registry::test_registration_module_inventory_is_complete AssertionError: assert {'flashinfer....cat_ops', ...} == {'flashinfer....cat_ops', ...} Opened 1 month ago | nvamyt | closed - completed | bug needs-triage ci: health | 1 | 0 | 1 day ago |
#5145 [Bug] JIT sampler built with CUDA 13 is reused by CUDA 12.9 Ray workers and fails at runtime Opened 1 day ago | ZFXzzz | open | needs-triage | 2 | 0 | 1 day ago |