Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-08 | 1 | 0 | 0 | 0 | 5 |
| 2026-09-07 | 2 | 3 | 0 | 0 | 2 |
| 2026-09-06 | 2 | 1 | 0 | 0 | 6 |
| 2026-09-05 | 1 | 1 | 0 | 0 | 10 |
| 2026-09-04 | 4 | 3 | 0 | 0 | 7 |
| 2026-09-03 | 5 | 1 | 0 | 0 | 118 |
| 2026-09-02 | 11 | 4 | 0 | 0 | 0 |
| 2026-09-01 | 2 | 1 | 0 | 0 | 0 |
| 2026-08-31 | 2 | 1 | 0 | 0 | 0 |
| 2026-08-30 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-29 | 4 | 0 | 0 | 0 | 0 |
| 2026-08-28 | 2 | 1 | 0 | 0 | 0 |
| 2026-08-27 | 2 | 1 | 0 | 0 | 0 |
| 2026-08-26 | 2 | 1 | 0 | 0 | 0 |
Opened: 26
Closed: 13
Comments: 0
Events: 0
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#7103 AsyncGRPO FSDP2 weight sync hangs: interleaved full_tensor() and is_checkpoint_format=True Opened 2 hours ago | sfc-gh-truwase | open | No labels | 0 | 0 | 2 hours ago |
#5191 Align trl skills CLI DX with HF CLI conventions Opened 6 months ago | albertvillanova | closed - not_planned | No labels | 3 | 0 | 8 hours ago |
#7100 `num_tokens` overcounted under TP Opened 9 hours ago | qgallouedec | open | No labels | 0 | 0 | 9 hours ago |
#7093 CI emits 462 `transformers` `Trainer` log warnings: "The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config" Opened 17 hours ago | albertvillanova | open | No labels | 4 | 0 | 12 hours ago |
#7086 RLOO uses higher-variance KL Estimator and should use lower variance k3 Opened 1 day ago | AMindToThink | open | No labels | 1 | 0 | 16 hours ago |
#5127 MGPO Opened 7 months ago | damoonsh | closed - not_planned | No labels | 2 | 0 | 1 day ago |
#5025 Maximum Likelihood Reinforcement Learning Opened 7 months ago | catherinelee274 | closed - not_planned | No labels | 1 | 0 | 1 day ago |
#3549 Packing sequences for memory efficiency in GRPO and other preference learning implementations Opened 1 year ago | RameshArvind | open | β¨ enhancement π GRPO | 11 | 1 | 1 day ago |
#7071 SFTTrainer treats an empty config._name_or_path as a Hub repo id Opened 3 days ago | victorzhong0110 | closed - not_planned | No labels | 3 | 0 | 2 days ago |
#7083 Support precomputed reference log-probs in fused DPO training Opened 2 days ago | alay2shah | open | No labels | 0 | 0 | 2 days ago |
#6833 RFC: Add single-image vision-language model support to A2POTrainer Opened 19 days ago | DimensionSTP | open | No labels | 1 | 1 | 2 days ago |
#6832 RFC: Add end-to-end vision-language model support to SDPOTrainer Opened 19 days ago | DimensionSTP | open | No labels | 1 | 0 | 2 days ago |
#6831 RFC: Support A2PO with DeepSpeed ZeRO-3 via an explicit reference model Opened 19 days ago | DimensionSTP | open | No labels | 1 | 0 | 2 days ago |
#4339 Packing with VLMs Opened 11 months ago | jiosephlee | open | β¨ enhancement π SFT | 5 | 0 | 3 days ago |
#7063 [Tracking] Move off liger-kernel and make the memory-efficient loss the default Opened 3 days ago | qgallouedec | open | No labels | 1 | 0 | 3 days ago |
#6970 Bug: pack_dataset BFD fails when all sequences are empty Opened 10 days ago | YZJF | closed - not_planned | No labels | 3 | 0 | 3 days ago |
#4184 [Feature] Add a bool enable/disable flag for `entropy logging` Opened 11 months ago | steveepreston | closed - not_planned | β¨ enhancement π SFT | 9 | 0 | 3 days ago |
#6221 [Feature] Reasoning reward utilities: group-relative length-scaled accuracy (GRPO-LEAD) + graduated format reward Opened 2 months ago | surajsharan | open | No labels | 4 | 0 | 3 days ago |
#7012 PPOTrainer statistics average per-micro-batch means without weighting by micro-batch size Opened 6 days ago | behroozazarkhalili | closed - not_planned | No labels | 1 | 0 | 3 days ago |
#7028 Does hf jobs uv run --image huggingface/trl actually use the preinstalled TRL? Opened 5 days ago | albertvillanova | closed - completed | No labels | 4 | 0 | 4 days ago |
#7024 MiniLLM: right-padding leaks into the reverse-KL advantage, and length normalization depends on batch padding Opened 5 days ago | behroozazarkhalili | open | No labels | 5 | 0 | 4 days ago |
#7047 GRPO loss-path metrics average micro-batch entries, so uneven accumulation windows are weighted by their size Opened 4 days ago | behroozazarkhalili | open | π bug | 6 | 0 | 4 days ago |
#6961 CI emits `transformers` warning log: Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'} Opened 11 days ago | albertvillanova | closed - completed | No labels | 4 | 0 | 4 days ago |
#6857 Tiny test models are incompatible with Flash Attention (head_size=2) Opened 18 days ago | albertvillanova | open | No labels | 1 | 1 | 4 days ago |
#6877 Bug: unpair IterableDatasetDict preference datasets Opened 16 days ago | YZJF | open | No labels | 1 | 0 | 4 days ago |