Repository Issue Activity (beta)

NVIDIA/Megatron-LM

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
393
New in 7 Days
21
Closed in 7 Days
11
Average Open Age
163 days
Stale 30+ Days
239
Stale 90+ Days
143
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-2400000
2026-08-2300265
2026-08-2201034
2026-08-212463314
2026-08-20633167
2026-08-19430137
2026-08-187023014
2026-08-17205134
2026-08-16006266
2026-08-151022411
2026-08-14200219
2026-08-131452112
2026-08-12321149
2026-08-111232513
This Week

Opened: 19

Closed: 11

Comments: 13

Events: 101

Top Labels
community-request (523)
bug (432)
stale (301)
enhancement (235)
question (184)
module: moe (97)
waiting-on-customer (79)
waiting-on-maintainers (63)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#6501 [feat] GTP+DCP

Opened 11 days ago
fanshiqing
closed - completed
nemotron
0021 hours ago

#1784 [ENHANCEMENT] MoE support in report_theoretical_memory

Opened 1 year ago
JungHoyoun
open
enhancement
module: moe
community-request
111 day ago

#6747 [VPP] Defer embedding initialization sync until all local model chunks are built

Opened 3 days ago
SonglinLife
open
bug
community-request
waiting-on-maintainers
001 day ago

#6757 [ROADMAP][2026 Q3] Megatron Core MoE Roadmap

Opened 3 days ago
buptzyb
open
call for contribution
101 day ago

#6729 triton KV append kernel hard-asserts CUDA, blocking non-CUDA Triton backends

Opened 3 days ago
tengqm
open
community-request
001 day ago

#6708 GRPO crashes with --transformer-impl local when RL training CUDA graphs are disabled

Opened 4 days ago
tengqm
open
community-request
002 days ago

#6540 [feat] GTP+Symmetric Memory Registration buffer

Opened 10 days ago
fanshiqing
closed - completed
nemotron
002 days ago

#6700 [RL][Performance] Avoid full-vocabulary TP gather for selected-token logprobs

Opened 4 days ago
chengcuiping
open
enhancement
community-request
waiting-on-maintainers
002 days ago

#1455 [QUESTION]Is there any plan to make custom_fsdp compatible with PP?

Opened 1 year ago
XCD4P
open
question
community-request
waiting-on-maintainers
203 days ago

#6600 MFSDP v2: Support EP composability with grouped 3D expert weights

Opened 6 days ago
wujingyue
open
No labels
203 days ago

#6585 [BUG] pretrain_vlm.py silently ignores --rotary-base: LLaVAModel always uses the default 10000

Opened 7 days ago
adityaghai07
closed - completed
community-request
003 days ago

#6087 [MFSDP V2] Post-wrap load_state_dict(assign=True) does not resume from the loaded checkpoint

Opened 27 days ago
chengcuiping
closed - completed
community-request
1503 days ago

#4815 [ROADMAP][2026 Q2] Megatron Core MoE Roadmap

Opened 3 months ago
Victarry
closed - completed
community-request
call for contribution
493 days ago

#6656 [BUG] FP32 backward fails in VocabParallelCrossEntropy with custom-autograd view/in-place error

Opened 5 days ago
fwerkor
open
community-request
003 days ago

#6363 [feat] GTP+2D Mxfp8

Opened 16 days ago
fanshiqing
open
nemotron
003 days ago

#6660 [Megatron-FSDP] Gradients doubled every microbatch (2^N amplification) when a grad buffer's DP group has size 1 (EP=DP or DP=1) with optim_grads_params

Opened 5 days ago
XiongFenghhh
closed - not_planned
bug
community-request
403 days ago

#6491 Offload MoE expert weights to pinned host memory to enable training at smaller EP

Opened 12 days ago
FFGGSSJJ
open
enhancement
community-request
waiting-on-customer
1103 days ago

#6720 [GTP][Muon] Optimize Muon step performance

Opened 4 days ago
wanyingw
open
No labels
003 days ago

#6392 GLM-5.2 training support tracking

Opened 14 days ago
buptzyb
open
No labels
103 days ago

#6714 MFSDP v2: overlap DP-inner and DP-outer communication

Opened 4 days ago
wujingyue
open
module: megatron-fsdp
MFSDPv2
103 days ago

#6111 MoE router padding mask fails with expert bias during training

Opened 26 days ago
cuichenx
closed - completed
No labels
003 days ago

#5755 [BUG] TransformerConfig accepts num_attention_heads not divisible by num_query_groups, then crashes with a cryptic shape error inside attention

Opened 1 month ago
huthvincent
open
bug
community-request
203 days ago

#3083 Profiling args are missing on the main brench

Opened 7 months ago
Skylion007
open
bug
community-request
waiting-on-customer
203 days ago

#6719 Pipeline the layer-sharded Muon all_to_all to hide exchange latency (and cut peak memory)

Opened 4 days ago
wanyingw
open
enhancement
004 days ago

#5787 [BUG] `ParamAndGradBuffer.update_main_grads` repeats the global-shape view for cached main gradients (rowwise-TP silent-wrong / deadlock)

Opened 1 month ago
yuhezhang-ai
closed - not_planned
bug
module: megatron-fsdp
204 days ago

Rows per page:

1–25 of 1,581