Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-24 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-23 | 0 | 0 | 2 | 6 | 5 |
| 2026-08-22 | 0 | 1 | 0 | 3 | 4 |
| 2026-08-21 | 2 | 4 | 6 | 33 | 14 |
| 2026-08-20 | 6 | 3 | 3 | 16 | 7 |
| 2026-08-19 | 4 | 3 | 0 | 13 | 7 |
| 2026-08-18 | 7 | 0 | 2 | 30 | 14 |
| 2026-08-17 | 2 | 0 | 5 | 13 | 4 |
| 2026-08-16 | 0 | 0 | 6 | 26 | 6 |
| 2026-08-15 | 1 | 0 | 2 | 24 | 11 |
| 2026-08-14 | 2 | 0 | 0 | 21 | 9 |
| 2026-08-13 | 1 | 4 | 5 | 21 | 12 |
| 2026-08-12 | 3 | 2 | 1 | 14 | 9 |
| 2026-08-11 | 1 | 2 | 3 | 25 | 13 |
Opened: 19
Closed: 11
Comments: 13
Events: 101
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#6501 [feat] GTP+DCP Opened 11 days ago | fanshiqing | closed - completed | nemotron | 0 | 0 | 21 hours ago |
#1784 [ENHANCEMENT] MoE support in report_theoretical_memory Opened 1 year ago | JungHoyoun | open | enhancement module: moe community-request | 1 | 1 | 1 day ago |
#6747 [VPP] Defer embedding initialization sync until all local model chunks are built Opened 3 days ago | SonglinLife | open | bug community-request waiting-on-maintainers | 0 | 0 | 1 day ago |
#6757 [ROADMAP][2026 Q3] Megatron Core MoE Roadmap Opened 3 days ago | buptzyb | open | call for contribution | 1 | 0 | 1 day ago |
#6729 triton KV append kernel hard-asserts CUDA, blocking non-CUDA Triton backends Opened 3 days ago | tengqm | open | community-request | 0 | 0 | 1 day ago |
#6708 GRPO crashes with --transformer-impl local when RL training CUDA graphs are disabled Opened 4 days ago | tengqm | open | community-request | 0 | 0 | 2 days ago |
#6540 [feat] GTP+Symmetric Memory Registration buffer Opened 10 days ago | fanshiqing | closed - completed | nemotron | 0 | 0 | 2 days ago |
#6700 [RL][Performance] Avoid full-vocabulary TP gather for selected-token logprobs Opened 4 days ago | chengcuiping | open | enhancement community-request waiting-on-maintainers | 0 | 0 | 2 days ago |
#1455 [QUESTION]Is there any plan to make custom_fsdp compatible with PP? Opened 1 year ago | XCD4P | open | question community-request waiting-on-maintainers | 2 | 0 | 3 days ago |
#6600 MFSDP v2: Support EP composability with grouped 3D expert weights Opened 6 days ago | wujingyue | open | No labels | 2 | 0 | 3 days ago |
#6585 [BUG] pretrain_vlm.py silently ignores --rotary-base: LLaVAModel always uses the default 10000 Opened 7 days ago | adityaghai07 | closed - completed | community-request | 0 | 0 | 3 days ago |
#6087 [MFSDP V2] Post-wrap load_state_dict(assign=True) does not resume from the loaded checkpoint Opened 27 days ago | chengcuiping | closed - completed | community-request | 15 | 0 | 3 days ago |
#4815 [ROADMAP][2026 Q2] Megatron Core MoE Roadmap Opened 3 months ago | Victarry | closed - completed | community-request call for contribution | 4 | 9 | 3 days ago |
#6656 [BUG] FP32 backward fails in VocabParallelCrossEntropy with custom-autograd view/in-place error Opened 5 days ago | fwerkor | open | community-request | 0 | 0 | 3 days ago |
#6363 [feat] GTP+2D Mxfp8 Opened 16 days ago | fanshiqing | open | nemotron | 0 | 0 | 3 days ago |
#6660 [Megatron-FSDP] Gradients doubled every microbatch (2^N amplification) when a grad buffer's DP group has size 1 (EP=DP or DP=1) with optim_grads_params Opened 5 days ago | XiongFenghhh | closed - not_planned | bug community-request | 4 | 0 | 3 days ago |
#6491 Offload MoE expert weights to pinned host memory to enable training at smaller EP Opened 12 days ago | FFGGSSJJ | open | enhancement community-request waiting-on-customer | 11 | 0 | 3 days ago |
#6720 [GTP][Muon] Optimize Muon step performance Opened 4 days ago | wanyingw | open | No labels | 0 | 0 | 3 days ago |
#6392 GLM-5.2 training support tracking Opened 14 days ago | buptzyb | open | No labels | 1 | 0 | 3 days ago |
#6714 MFSDP v2: overlap DP-inner and DP-outer communication Opened 4 days ago | wujingyue | open | module: megatron-fsdp MFSDPv2 | 1 | 0 | 3 days ago |
#6111 MoE router padding mask fails with expert bias during training Opened 26 days ago | cuichenx | closed - completed | No labels | 0 | 0 | 3 days ago |
#5755 [BUG] TransformerConfig accepts num_attention_heads not divisible by num_query_groups, then crashes with a cryptic shape error inside attention Opened 1 month ago | huthvincent | open | bug community-request | 2 | 0 | 3 days ago |
#3083 Profiling args are missing on the main brench Opened 7 months ago | Skylion007 | open | bug community-request waiting-on-customer | 2 | 0 | 3 days ago |
#6719 Pipeline the layer-sharded Muon all_to_all to hide exchange latency (and cut peak memory) Opened 4 days ago | wanyingw | open | enhancement | 0 | 0 | 4 days ago |
#5787 [BUG] `ParamAndGradBuffer.update_main_grads` repeats the global-shape view for cached main gradients (rowwise-TP silent-wrong / deadlock) Opened 1 month ago | yuhezhang-ai | closed - not_planned | bug module: megatron-fsdp | 2 | 0 | 4 days ago |