Repository Issue Activity (beta)

nvidia/transformerengine

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
163
New in 7 Days
4
Closed in 7 Days
0
Average Open Age
295 days
Stale 30+ Days
133
Stale 90+ Days
93
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-2000444
2026-09-1910624163
2026-09-1800000
2026-09-1700000
2026-09-1610000
2026-09-1520000
2026-09-1400000
2026-09-1300000
2026-09-1200000
2026-09-1100000
2026-09-1002000
2026-09-0901000
2026-09-0800000
2026-09-0700000
This Week

Opened: 4

Closed: 0

Comments: 10

Events: 28

Top Labels
bug (190)
question (80)
attention (52)
enhancement (42)
MoE (36)
build (33)
performance (14)
megatron (13)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#2541 no attention backend support for softmax = learnable and cp_comm_type = p2p

Opened 9 months ago
jordane95
open
bug
3112 hours ago

#2986 Expose Batch Invariant Kernels

Opened 4 months ago
wdykas
open
enhancement
2015 hours ago

#1761 Create option to control data type for tensor parallel all-reduce

Opened 1 year ago
veritas9872
open
enhancement
5015 hours ago

#2995 Reduce CPU overheads in te Sequential GroupedLinear Op

Opened 4 months ago
vthumbe1503
open
No labels
4015 hours ago

#2405 CUDA extension loading should be modular

Opened 10 months ago
sbhavani
open
bug
111 day ago

#3437 Tied weights do not accumulate with delayed wgrad

Opened 23 days ago
wujingyue
open
No labels
201 day ago

#3550 Preserve NaN through half-precision MXFP8 amax reductions

Opened 1 day ago
sylvesterkaczmarek
open
No labels
001 day ago

#789 Request for Adaptive Layer Norm MLP

Opened 2 years ago
fordflip
open
enhancement
903 days ago

#3528 FA4 returns an all-zero output for causal attention on SM100 after flash-attention #2490

Opened 5 days ago
nvegesna-netizen
open
No labels
303 days ago

#3518 MXFP8 GEMM rejects compact unswizzled scales that quantize_ accepts

Opened 5 days ago
wujingyue
open
No labels
205 days ago

#3520 Graph-Safe FP8 BlockScaling Support on Blackwell for GroupedLinear

Opened 5 days ago
vthumbe1503
open
No labels
005 days ago

#3487 [PyTorch][Attention] THD P2P context-parallel regression when padded cu_seqlens are value-equal but not object-identical

Opened 15 days ago
cuichenx
open
attention
006 days ago

#2807 Adaptive Compression

Opened 6 months ago
guanchenl
open
question
426 days ago

#3312 thd + attention_dropout routes to the composite cuDNN SDPA engine and costs 5x on Blackwell

Opened 2 months ago
bzantium
open
attention
0010 days ago

#1688 fp8_model_init does nothing when used with FSDP2

Opened 1 year ago
MaciejBalaNV
closed - completed
bug
2010 days ago

#1436 If no Windows support is planned for 5090 as it wasn't for 4090, pass along to corporate to mention this in advertisement

Opened 2 years ago
NeedsMoar
closed - not_planned
question
5010 days ago

#1532 How can we use te.Linear with weight parallel?

Opened 2 years ago
zigzagcai
open
question
3010 days ago

#2703 Performance bottleneck in TransformerLayer and Attention when using attention mask

Opened 7 months ago
yang-cx
open
bug
1010 days ago

#3484 [Bug] NaN expert weight gradients at num_groups == 1 with a padded token buffer for SReLU

Opened 16 days ago
GarlGuo
closed - completed
bug
1012 days ago

#3028 [PyTorch] Integrate cuDNN GQA + DSA backend into DotProductAttention

Opened 4 months ago
nvMelissa
open
attention
0112 days ago

#3481 [Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it

Opened 16 days ago
yuweih205
open
attention
4012 days ago

#3249 [Bug] FusedAttention THD + learnable softmax backward IMA (cuDNN err 700) on Hopper

Opened 2 months ago
fy1214
open
bug
attention
6012 days ago

#3248 FP8 DPA future token leakage

Opened 2 months ago
jeromeku
open
bug
attention
1012 days ago

#3218 cuDNN to accept actual seqlens in addition to cumulative seqlens

Opened 2 months ago
sudhakarsingh27
open
attention
1012 days ago

#3188 Plumb dbias request in isolation into the common backend selector

Opened 2 months ago
KshitijLakhani
open
bug
attention
0012 days ago

Rows per page:

1–25 of 598