Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-24 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-23 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-22 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-21 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-20 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-19 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-18 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-17 | 0 | 0 | 0 | 0 | 19 |
| 2026-08-16 | 24 | 0 | 0 | 0 | 5 |
| 2026-08-15 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-14 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-13 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-12 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-11 | 0 | 0 | 0 | 0 | 0 |
Opened: 0
Closed: 0
Comments: 0
Events: 0
No label distribution is available yet.
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#299 GPQA debug mode ignores an explicit --examples sample count Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#298 Eval CLI --examples crashes non-debug AIME and GPQA because n_repeats stays 8 Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#296 Browser HTML conversion can leak global html2text monkey patches after an exception Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#294 AIME and GPQA treat num_examples=0 as the full dataset Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#292 HealthBench debug mode ignores an explicit --examples override Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#290 Distributed chat executes each tool call independently on every model-parallel rank Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#288 Eval progress helper crashes on an empty sample set Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#286 Local Harmony tokenizer assigns token ID 200018 to two special tokens Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#283 Nullable Responses request fields can crash the local server Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#281 Responses stub backend does not reset fake output on new requests Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#279 Responses Ollama backend injects token 0 when no streamed token is ready Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#277 Responses API default checkpoint path is not expanded Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#275 Responses vLLM backend passes TP environment value as a string Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#273 Responses API Triton backend binds CUDA devices using global rank Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#271 Responses API Triton backend creates KV caches on CPU Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#270 Distributed Torch and Triton sampling can diverge across model-parallel ranks Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#268 Text generation default unlimited limit passes an invalid max_tokens value Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#266 Distributed Torch inference binds CUDA devices using global rank Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#264 Chat Completions reasoning is injected into eval prompt history Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#262 Eval samplers retry permanent API errors forever Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#260 HealthBench bootstrap uncertainty is not reproducible Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#257 Merged eval summary mislabels HealthBench subset names Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#256 HealthBench usage metadata crashes with the chat_completions sampler Opened 8 days ago | sylvesterkaczmarek | open | No labels | 0 | 0 | 8 days ago |
#254 ABCD eval grader can score an earlier answer instead of the model's final answer Opened 9 days ago | sylvesterkaczmarek | open | No labels | 0 | 1 | 9 days ago |
#240 MCP mutation probe - please ignore Opened 6 months ago | oaimcpatlas-alt | open | No labels | 1 | 0 | 3 months ago |