Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-24 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-23 | 0 | 0 | 0 | 0 | 4 |
| 2026-08-22 | 2 | 0 | 1 | 3 | 3 |
| 2026-08-21 | 3 | 3 | 1 | 4 | 1 |
| 2026-08-20 | 1 | 2 | 2 | 2 | 3 |
| 2026-08-19 | 2 | 0 | 3 | 5 | 4 |
| 2026-08-18 | 1 | 0 | 0 | 0 | 2 |
| 2026-08-17 | 0 | 0 | 3 | 3 | 2 |
| 2026-08-16 | 2 | 0 | 3 | 7 | 5 |
| 2026-08-15 | 1 | 0 | 0 | 0 | 1 |
| 2026-08-14 | 0 | 0 | 1 | 3 | 1 |
| 2026-08-13 | 0 | 0 | 1 | 1 | 1 |
| 2026-08-12 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-11 | 0 | 0 | 0 | 0 | 1 |
Opened: 9
Closed: 5
Comments: 7
Events: 14
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#4038 minerva_math/putnam_axiom/leaderboard-math: thousands-separator comma strip fuses bare digit tuples ("0,1" -> "01") Opened 2 days ago | feiiiiii5 | open | No labels | 0 | 0 | 2 days ago |
#4036 minerva_math/putnam_axiom/leaderboard-math: sqrt shorthand normalization corrupts indexed roots \sqrt[3]{x} Opened 2 days ago | feiiiiii5 | open | No labels | 0 | 0 | 2 days ago |
#4028 --cache_requests delete fails when the cache directory does not exist Opened 3 days ago | sdivyanshu90 | open | No labels | 1 | 0 | 2 days ago |
#3046 Cache fails when repeats are greater than 1 Opened 1 year ago | eyuansu62 | open | bug | 1 | 1 | 2 days ago |
#3749 Feature request: canonical reproducibility-bundle export in scripts/ Opened 4 months ago | topeuph-ai | open | No labels | 12 | 0 | 2 days ago |
#4031 minerva_math: REMOVED_EXPRESSIONS entry "ft" corrupts \left (substring replace inside control sequences) Opened 3 days ago | rcfa | open | No labels | 0 | 0 | 3 days ago |
#4030 minerva_math: identical answers scored wrong when sympy cannot parse them (tuples, intervals, matrices) Opened 3 days ago | rcfa | open | No labels | 0 | 0 | 3 days ago |
#4002 Scalar `seed` in a YAML config crashes with TypeError; `seed: 0` silently disables seeding Opened 9 days ago | winklemad | closed - completed | No labels | 0 | 0 | 3 days ago |
#4011 `_parse_dict_args` is a no-op under postponed annotations, and `DICT_KEYS` is unused Opened 6 days ago | winklemad | closed - completed | No labels | 0 | 0 | 3 days ago |
#3962 [Proposal] OpenEval Import/Export Support Opened 26 days ago | adhabnr-ux | closed - not_planned | No labels | 2 | 0 | 3 days ago |
#3558 think_end_token support needed for local-chat-completions mode Opened 7 months ago | kediaharshit9 | closed - completed | No labels | 1 | 2 | 4 days ago |
#3460 max_model_length for gemma3-12b is set to 2048 (_DEFAULT_MAX_LENGTH) with the hf processor Opened 9 months ago | bodasadallah | closed - completed | No labels | 0 | 0 | 4 days ago |
#4022 Feature request: EvalPort import/export for per-document samples and results Opened 4 days ago | adhabnr-ux | open | No labels | 0 | 0 | 4 days ago |
#4019 Proposal: Wilson CIs + content-hash integrity field on results Opened 5 days ago | CSOAI-ORG | open | No labels | 1 | 0 | 5 days ago |
#3017 Support for using a remote /tokenize API endpoint as the tokenizer Opened 1 year ago | furkancoskun | open | feature request | 2 | 4 | 5 days ago |
#4017 Proposal: Wilson CI bounds, MDE, and optional BH-FDR flags on results output Opened 5 days ago | CSOAI-ORG | open | No labels | 0 | 0 | 5 days ago |
#4007 Generative-task scorers cannot distinguish an unparseable response from a wrong answer (census: 2,936/4,524 tasks affected) Opened 8 days ago | Rajveer-code | open | No labels | 2 | 0 | 7 days ago |
#4005 --use_cache crashes before inference for multimodal image requests Opened 8 days ago | sdivyanshu90 | open | No labels | 0 | 0 | 8 days ago |
#3080 Fix Metric Calculation for Repeats Opened 1 year ago | baberabb | open | No labels | 6 | 0 | 8 days ago |
#3493 Standardizing Task Formats Opened 7 months ago | baberabb | open | No labels | 1 | 0 | 9 days ago |
#3881 Request cache key ignores generation_kwargs -> silently reuses cached instances across different sampling parameters Opened 2 months ago | AUTHENSOR | open | No labels | 1 | 0 | 10 days ago |
#3498 what's the proper config for aime24 evalated on Qwen3? Opened 7 months ago | Lynnzake | open | No labels | 2 | 2 | 11 days ago |
#3989 Fix broken Discord link in contributing guide Opened 14 days ago | sdivyanshu90 | open | No labels | 0 | 0 | 14 days ago |
#3930 [Bug] afrixnli prompt_1: doc_to_text uses .format() braces, so premise and hypothesis are never substituted (35 tasks) Opened 1 month ago | AmosBunde | closed - completed | No labels | 2 | 0 | 14 days ago |
#3859 SCROLLS tasks fail to load with recent datasets (dataset scripts no longer supported) Opened 2 months ago | bongho | closed - completed | No labels | 2 | 0 | 15 days ago |