Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-12 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-11 | 0 | 0 | 0 | 0 | 77 |
| 2026-09-10 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-09 | 1 | 2 | 0 | 0 | 0 |
| 2026-09-08 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-07 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-06 | 1 | 2 | 0 | 0 | 0 |
| 2026-09-05 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-04 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-03 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-02 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-01 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-31 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-30 | 0 | 0 | 0 | 0 | 0 |
Opened: 4
Closed: 4
Comments: 0
Events: 0
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#920 reset_batch() wipes RejectionSamplingState every batch, so "episode" mode never accumulates across batches Opened 2 days ago | AmirF194 | open | No labels | 0 | 0 | 2 days ago |
#916 Bug: `reinforce_plus_plus_baseline` silently produces zero advantages when `rollout.n = 1` Opened 3 days ago | yanan1116 | open | bug | 1 | 0 | 3 days ago |
#432 context compression Opened 6 months ago | klsr27 | closed - not_planned | question | 0 | 0 | 4 days ago |
#391 The actor/entropy loss appears somewhat abnormal—it's too large. We launched DeepSWE training based on the Qwen3-4B model, using the following script: Opened 7 months ago | Cesilina | closed - not_planned | No labels | 2 | 0 | 4 days ago |
#915 check_correctness / lcb_check_correctness_v2 leave a zombie process after killing a timed-out test run Opened 6 days ago | AmirF194 | open | No labels | 1 | 0 | 4 days ago |
#914 R1ToolParser.tool_output_end doesn't match DeepSeek's own chat_template tool-output token Opened 7 days ago | AmirF194 | open | No labels | 1 | 0 | 6 days ago |
#782 Fireworks loss normalization depends on fwd/bwd chunking after #781 Opened 2 months ago | signalrush | closed - completed | No labels | 2 | 0 | 6 days ago |
#829 Dataset row ids containing ':' silently collapse every rollout into one GRPO group Opened 1 month ago | LEE-CHENYU | closed - completed | No labels | 2 | 0 | 7 days ago |
#901 No train/held-out split for claw_eval and skillsbench Opened 19 days ago | yanan1116 | open | No labels | 3 | 0 | 8 days ago |
#468 Sandboxed code execution for RL rollouts Opened 5 months ago | congwang-mk | open | No labels | 2 | 0 | 13 days ago |
#605 AgentWorkflowPPOTrainer: GRPO groups by trajectory.uid instead of prompt/agent key Opened 3 months ago | TtLuckyyy | open | question | 1 | 0 | 16 days ago |
#897 async_save=true on the default FSDP strategy never writes the checkpoint tracker file, so resume_mode:auto silently restarts from scratch Opened 22 days ago | AmirF194 | open | No labels | 2 | 0 | 17 days ago |
#816 OpenAIEngine.completion() retokenization fallback can desync completion_ids from logprobs Opened 1 month ago | AmirF194 | open | No labels | 0 | 0 | 1 month ago |
#655 Dockerfile RUN-replay corrupts Harbor task setup on modal/daytona (line-continuation mangling + double-apply on prebuilt images) Opened 3 months ago | signalrush | closed - completed | No labels | 0 | 0 | 1 month ago |
#388 There may be some compatibility conflicts between Transformers 5.0 and older vLLM. Opened 7 months ago | SyncLionPaw | closed - not_planned | No labels | 0 | 0 | 1 month ago |
#380 Supporting deep agents and/or automatically generate sub-agents without predefined flow Opened 8 months ago | Wonder1905 | closed - not_planned | No labels | 2 | 0 | 1 month ago |
#354 actor_rollout_ref.actor.loss_agg_mode=seq-mean-token-sum in examples/swe/train_deepswe_32b.sh Opened 9 months ago | p81sunshine | closed - not_planned | No labels | 2 | 0 | 2 months ago |
#717 rllm train crashes at first validation: 'FixedEvaluatorHooks' object has no attribute 'warm_queue' Opened 2 months ago | Chanbinski2 | open | No labels | 1 | 0 | 2 months ago |
#378 Kind Kubernetes settings Opened 8 months ago | bluesky983 | closed - not_planned | No labels | 0 | 0 | 2 months ago |
#350 ppo_mini_batch_size for Non-Cumulative Agents Opened 9 months ago | huseyinatahaninan | closed - not_planned | question | 2 | 0 | 2 months ago |
#712 rllm sft crashes with FileNotFoundError — SFT backend config YAMLs excluded from built package Opened 2 months ago | Chanbinski2 | closed - completed | No labels | 0 | 0 | 2 months ago |
#704 How to run PPO for agents? Opened 3 months ago | miaoyuchun | open | question | 1 | 0 | 3 months ago |
#355 How to implement trajectory-level masking for "invalid turns" (SimpleTIR) in rllm? Opened 9 months ago | BabelTower | closed - not_planned | No labels | 1 | 0 | 3 months ago |
#226 DeepScaleR 8k reproduction collapses after ~300 steps with sequence length explosion Opened 1 year ago | zhqwqwq | closed - not_planned | No labels | 6 | 0 | 3 months ago |
#434 Qwen3.5 support Opened 6 months ago | HaishuoFang | open | question | 13 | 0 | 3 months ago |