Repository Issue Activity (beta)

huggingface/trl

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
124
New in 7 Days
28
Closed in 7 Days
14
Average Open Age
112 days
Stale 30+ Days
38
Stale 90+ Days
10
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-0810005
2026-09-0723002
2026-09-0621006
2026-09-05110010
2026-09-0443007
2026-09-035100118
2026-09-02114000
2026-09-0121000
2026-08-3121000
2026-08-3000000
2026-08-2940000
2026-08-2821000
2026-08-2721000
2026-08-2621000
This Week

Opened: 26

Closed: 13

Comments: 0

Events: 0

Top Labels
πŸ› bug (537)
πŸ‹ GRPO (374)
✨ enhancement (242)
❓ question (171)
πŸ‹ SFT (135)
⚑ PEFT (125)
⚑accelerate (98)
πŸ‹ DPO (95)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#7103 AsyncGRPO FSDP2 weight sync hangs: interleaved full_tensor() and is_checkpoint_format=True

Opened 2 hours ago
sfc-gh-truwase
open
No labels
002 hours ago

#5191 Align trl skills CLI DX with HF CLI conventions

Opened 6 months ago
albertvillanova
closed - not_planned
No labels
308 hours ago

#7100 `num_tokens` overcounted under TP

Opened 9 hours ago
qgallouedec
open
No labels
009 hours ago

#7093 CI emits 462 `transformers` `Trainer` log warnings: "The tokenizer has new PAD/BOS/EOS tokens that differ from the model config and generation config"

Opened 17 hours ago
albertvillanova
open
No labels
4012 hours ago

#7086 RLOO uses higher-variance KL Estimator and should use lower variance k3

Opened 1 day ago
AMindToThink
open
No labels
1016 hours ago

#5127 MGPO

Opened 7 months ago
damoonsh
closed - not_planned
No labels
201 day ago

#5025 Maximum Likelihood Reinforcement Learning

Opened 7 months ago
catherinelee274
closed - not_planned
No labels
101 day ago

#3549 Packing sequences for memory efficiency in GRPO and other preference learning implementations

Opened 1 year ago
RameshArvind
open
✨ enhancement
πŸ‹ GRPO
1111 day ago

#7071 SFTTrainer treats an empty config._name_or_path as a Hub repo id

Opened 3 days ago
victorzhong0110
closed - not_planned
No labels
302 days ago

#7083 Support precomputed reference log-probs in fused DPO training

Opened 2 days ago
alay2shah
open
No labels
002 days ago

#6833 RFC: Add single-image vision-language model support to A2POTrainer

Opened 19 days ago
DimensionSTP
open
No labels
112 days ago

#6832 RFC: Add end-to-end vision-language model support to SDPOTrainer

Opened 19 days ago
DimensionSTP
open
No labels
102 days ago

#6831 RFC: Support A2PO with DeepSpeed ZeRO-3 via an explicit reference model

Opened 19 days ago
DimensionSTP
open
No labels
102 days ago

#4339 Packing with VLMs

Opened 11 months ago
jiosephlee
open
✨ enhancement
πŸ‹ SFT
503 days ago

#7063 [Tracking] Move off liger-kernel and make the memory-efficient loss the default

Opened 3 days ago
qgallouedec
open
No labels
103 days ago

#6970 Bug: pack_dataset BFD fails when all sequences are empty

Opened 10 days ago
YZJF
closed - not_planned
No labels
303 days ago

#4184 [Feature] Add a bool enable/disable flag for `entropy logging`

Opened 11 months ago
steveepreston
closed - not_planned
✨ enhancement
πŸ‹ SFT
903 days ago

#6221 [Feature] Reasoning reward utilities: group-relative length-scaled accuracy (GRPO-LEAD) + graduated format reward

Opened 2 months ago
surajsharan
open
No labels
403 days ago

#7012 PPOTrainer statistics average per-micro-batch means without weighting by micro-batch size

Opened 6 days ago
behroozazarkhalili
closed - not_planned
No labels
103 days ago

#7028 Does hf jobs uv run --image huggingface/trl actually use the preinstalled TRL?

Opened 5 days ago
albertvillanova
closed - completed
No labels
404 days ago

#7024 MiniLLM: right-padding leaks into the reverse-KL advantage, and length normalization depends on batch padding

Opened 5 days ago
behroozazarkhalili
open
No labels
504 days ago

#7047 GRPO loss-path metrics average micro-batch entries, so uneven accumulation windows are weighted by their size

Opened 4 days ago
behroozazarkhalili
open
πŸ› bug
604 days ago

#6961 CI emits `transformers` warning log: Unrecognized keys in `rope_parameters` for 'rope_type'='default': {'full_attention', 'sliding_attention'}

Opened 11 days ago
albertvillanova
closed - completed
No labels
404 days ago

#6857 Tiny test models are incompatible with Flash Attention (head_size=2)

Opened 18 days ago
albertvillanova
open
No labels
114 days ago

#6877 Bug: unpair IterableDatasetDict preference datasets

Opened 16 days ago
YZJF
open
No labels
104 days ago

Rows per page:

1–25 of 1,657