Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-20 | 0 | 0 | 0 | 0 | 102 |
| 2026-09-19 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-18 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-17 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-16 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-15 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-14 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-13 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-12 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-11 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-10 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-09 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-08 | 2 | 1 | 0 | 0 | 0 |
| 2026-09-07 | 0 | 0 | 0 | 0 | 0 |
Opened: 1
Closed: 0
Comments: 0
Events: 0
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#499 [Bug] 权重同步未应用模型自带的 hf_to_vllm_mapper,wrapper 形态模型(如 model.language_model.* 前缀)在第一次权重更新即失败 Opened 5 days ago | TracebaK | open | No labels | 1 | 0 | 5 days ago |
#474 如何使用cot格式对Qwen3.5-0.8B进行全量微调 Opened 2 months ago | Peak925 | closed - completed | No labels | 0 | 0 | 13 days ago |
#496 求,什么时候支持一下RSI。 Opened 13 days ago | Peak925 | open | No labels | 0 | 0 | 13 days ago |
#495 BatchStratifiedSampler crashes with ZeroDivisionError when a domain_ratios entry has no matching rows Opened 13 days ago | AmirF194 | open | No labels | 0 | 0 | 13 days ago |
#423 Bug: image_grid_thw[image_index][0] IndexError: list index out of range Opened 5 months ago | luyouqi233 | closed - completed | No labels | 2 | 0 | 25 days ago |
#490 [RFC] Upstream Environment-Regularized Policy Optimization (ERPO) into ROLL Opened 25 days ago | adtureven | open | No labels | 0 | 0 | 25 days ago |
#485 Colocated full-parameter weight sync: CPU fallback costs ~17% of every training step on seccomp-managed containers; request a pidfd_getfd-free GPU transport Opened 1 month ago | Donec-x | open | No labels | 0 | 0 | 1 month ago |
#484 Megatron colocated LoRA weight sync fails on CUDA IPC and needs CPU staging fallback Opened 1 month ago | Su245811YZ | open | No labels | 0 | 0 | 1 month ago |
#482 Verify evals on Papers with Code Opened 1 month ago | NielsRogge | open | No labels | 0 | 0 | 1 month ago |
#475 [Bug] Advantage whitening fails with a single valid response token Opened 2 months ago | Jackie2049 | closed - completed | No labels | 0 | 0 | 2 months ago |
#329 Training hangs at actor-infer step with Qwen3-8B on an 8-GPU node Opened 8 months ago | UsernameFull | closed - completed | No labels | 6 | 0 | 2 months ago |
#316 AttributeError: 'NoneType' object has no attribute 'rename_key_' Opened 8 months ago | bastianxux | open | No labels | 2 | 0 | 2 months ago |
#411 Error: Qwen3.5-35B-A3B lora sft with mcore-adapter, error occurs when saving a ckpt. Opened 6 months ago | Unofish | open | No labels | 12 | 0 | 2 months ago |
#419 mcore-adapter从VLM的hfconfig中获取参数异常 Opened 5 months ago | YisuZhou | open | No labels | 3 | 0 | 2 months ago |
#150 deepspeed zero3 model_update error Opened 1 year ago | mst272 | open | No labels | 1 | 0 | 2 months ago |
#418 ValueError: There is no module or parameter named 'base_model' in Qwen2_5_VLForConditionalGeneration Opened 5 months ago | luyouqi233 | open | No labels | 6 | 0 | 2 months ago |
#442 LR scheduler progress can be inconsistent with dynamic batching in Megatron actor training Opened 4 months ago | 56546256576885 | open | No labels | 3 | 0 | 2 months ago |
#435 [Bug] Training/Inference crash with Qwen2/3-VL due to missing mm_token_type_ids in Collator Opened 5 months ago | Wangxiaoxiaoa | open | No labels | 1 | 0 | 2 months ago |
#205 间断显卡编号的设备映射问题(English translation: device_mapping cannot handle discontinuous GPU indexes) Opened 11 months ago | DFexpres | open | No labels | 3 | 0 | 2 months ago |
#279 weight update progress过于缓慢 Opened 10 months ago | luyouqi233 | open | No labels | 4 | 0 | 2 months ago |
#309 code_dapo agentic task fails Opened 9 months ago | canghongjian | open | No labels | 3 | 0 | 2 months ago |
#394 Very slow asynchronous GRPO training on 8×A100 Opened 6 months ago | 5SSjw | open | No labels | 4 | 0 | 2 months ago |
#230 ROLL 是否可以加速reward模型(非LLM)计算与LLM GRPO的高效协同训练 Opened 10 months ago | JoshonSmith | open | No labels | 1 | 0 | 2 months ago |
#407 LR scheduler exhausts early in agentic training with AgentNativeStepEnvManager Opened 6 months ago | shamanez | open | No labels | 5 | 0 | 2 months ago |
#472 [RFC] Add MindSpeed/Megatron Unit Test CI Pipeline Opened 3 months ago | UsernameFull | open | No labels | 0 | 0 | 3 months ago |