Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-24 | 0 | 0 | 1 | 3 | 0 |
| 2026-08-23 | 6 | 5 | 3 | 11 | 8 |
| 2026-08-22 | 0 | 0 | 0 | 1 | 1 |
| 2026-08-21 | 1 | 3 | 8 | 21 | 3 |
| 2026-08-20 | 2 | 0 | 5 | 8 | 3 |
| 2026-08-19 | 1 | 3 | 2 | 4 | 0 |
| 2026-08-18 | 2 | 1 | 2 | 17 | 4 |
| 2026-08-17 | 0 | 0 | 1 | 2 | 1 |
| 2026-08-16 | 3 | 1 | 0 | 1 | 6 |
| 2026-08-15 | 3 | 0 | 0 | 1 | 1 |
| 2026-08-14 | 0 | 1 | 0 | 0 | 0 |
| 2026-08-13 | 1 | 0 | 0 | 0 | 3 |
| 2026-08-12 | 0 | 0 | 0 | 0 | 1 |
| 2026-08-11 | 0 | 0 | 0 | 0 | 0 |
Opened: 12
Closed: 12
Comments: 21
Events: 65
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#11406 `tl.softmax` produces incorrect, non-row-wise results Opened 2 days ago | merlinsun | closed - completed | bug | 2 | 0 | 20 hours ago |
#11407 [Pipeliner] predicateOp drops the predicate on NotSpeculatable arithmetic Opened 2 days ago | alepot55 | closed - completed | No labels | 0 | 0 | 1 day ago |
#11328 [ClusterBarrierInsertion] Multicast TMA into a memdesc view misses its cluster barrier Opened 9 days ago | alepot55 | closed - completed | No labels | 7 | 0 | 1 day ago |
#11408 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | No labels | 0 | 0 | 1 day ago |
#11373 [SM80/A100] 2x perf regression 3.6.0 -> 3.7.1 at previously-optimal num_warps=4 (per-config optimum inverts); 13+ min compiles on large BLOCK sizes Opened 5 days ago | onesource2026 | open | No labels | 1 | 0 | 1 day ago |
#11393 OP ttg.local_atomic_scatter_rmw fails in ClusterBarrierAnalysis: LLVM ERROR: scratch buffer operations should not have any shared memory dependencies for Opened 3 days ago | whutsunxu | closed - completed | bug | 5 | 0 | 2 days ago |
#11405 [TritonGPU] ttg.warp_specialize capturing a tensor aborts with report_fatal_error and no diagnostic Opened 2 days ago | alepot55 | open | No labels | 0 | 0 | 2 days ago |
#11404 [Membar] A relaxed ttng.cluster_barrier is treated as a full cluster sync point Opened 2 days ago | alepot55 | open | No labels | 0 | 0 | 2 days ago |
#11402 `tl.condition` never terminates when the initial while condition is false Opened 2 days ago | merlinsun | open | bug | 0 | 0 | 2 days ago |
#11352 `tl.cat` fails with misleading errors when operands differ along `dim`, despite its assertion allowing it Opened 6 days ago | zhudada0120 | closed - completed | bug | 0 | 0 | 3 days ago |
#11037 [AMD/gfx950] tl.atomic_add inside a while loop hangs Opened 1 month ago | shuyaobi-afk | closed - completed | No labels | 4 | 0 | 3 days ago |
#11320 fp8 matmul on sm_120 runs at half the 8-bit rate Opened 10 days ago | dxqb | closed - completed | performance | 0 | 1 | 3 days ago |
#10907 [v.3.8.0] Release Tracker Opened 1 month ago | atalman | open | No labels | 23 | 0 | 3 days ago |
#10229 Triton NVIDIA backend: 'ttng.tensormap_create' op pipeliner doesn't know how to predicate this op (CUDA 13.0, sm_90 + sm_100) Opened 4 months ago | afierka-intel | closed - completed | No labels | 2 | 0 | 4 days ago |
#11378 [AMD] scf.for with runtime trip count miscompiled on gfx1151 Opened 4 days ago | neuhaus | open | bug | 0 | 0 | 4 days ago |
#11146 Loop body is generated 2^depth times when lowering nested loops Opened 23 days ago | teerthsharma | closed - completed | No labels | 1 | 0 | 5 days ago |
#10486 Probabilistic illegal memory access from a kernel with multiple tl.dot on H100 (sm_90) Opened 3 months ago | tarinduj | closed - completed | bug needs reproducer | 3 | 1 | 5 days ago |
#11366 [NVIDIA] Hopper wgmma IMA on Triton 3.7.1: Invalid __shared__ read (H20 / sm90) Opened 6 days ago | wyc-ruiker | closed - completed | No labels | 1 | 1 | 5 days ago |
#10987 Triton generates pathologically large PTX for chained fp32 division-by-broadcast feeding a reduction Opened 1 month ago | yushangdi | closed - completed | performance | 0 | 0 | 7 days ago |
#11327 AMD compilation crash in TritonAMDGPUPrepareIfCombining Opened 9 days ago | O1iveira7 | open | bug | 1 | 0 | 7 days ago |
#10331 Docs: bring-up recipe for Triton on NVIDIA Blackwell sm_121 (GB10 / DGX Spark / consumer Blackwell) Opened 3 months ago | eniktab | open | No labels | 6 | 0 | 7 days ago |
#11344 `desc.scatter()` emits `tile::scatter4` on consumer Blackwell (sm_12x) and fails in ptxas instead of erroring at compile time Opened 7 days ago | yashb98 | open | No labels | 0 | 0 | 7 days ago |
#9703 [Perf] TMA descriptor load followed by reduce regresses in 3.6.0 compared to 3.5.1 Opened 6 months ago | ngoyal2707 | open | performance | 6 | 0 | 8 days ago |
#11290 [tritongpu] Deduplicate the three hoistConvert*() driver functions in RemoveLayoutConversions Opened 11 days ago | aliharbaji | closed - completed | No labels | 1 | 0 | 8 days ago |
#11326 [Membar] No barrier between a caller's pending shared access and a callee's first one Opened 9 days ago | alepot55 | open | No labels | 0 | 0 | 9 days ago |