Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-07 | 0 | 0 | 0 | 0 | 1 |
| 2026-09-06 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-05 | 1 | 0 | 0 | 0 | 1 |
| 2026-09-04 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-03 | 0 | 0 | 0 | 0 | 240 |
| 2026-09-02 | 2 | 0 | 0 | 0 | 0 |
| 2026-09-01 | 1 | 0 | 0 | 0 | 0 |
| 2026-08-31 | 25 | 0 | 0 | 0 | 0 |
| 2026-08-30 | 3 | 0 | 0 | 0 | 0 |
| 2026-08-29 | 1 | 0 | 0 | 0 | 0 |
| 2026-08-28 | 2 | 0 | 0 | 0 | 0 |
| 2026-08-27 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-26 | 2 | 0 | 0 | 0 | 0 |
| 2026-08-25 | 1 | 0 | 0 | 0 | 0 |
Opened: 4
Closed: 0
Comments: 0
Events: 0
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#3216 fix(benchmarks): replace bare n_shots/n_problems asserts with explicit validation in remaining benchmarks Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 1 day ago |
#3245 ToolCorrectnessMetric reports type mismatches for identical calls in exact-match mode Opened 2 days ago | JuSe123456 | open | No labels | 0 | 0 | 2 days ago |
#3218 fix(scorer): quasi_contains_score silently mis-scores a bare string target Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 5 days ago |
#3210 Scorer.truth_identification_score returns percentages above 100 when predictions/targets contain duplicated indices Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 5 days ago |
#3202 get_turns_in_sliding_window silently yields empty windows when window_size <= 0 Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 5 days ago |
#3235 MMLU reuses the first task's few-shot examples across subjects Opened 5 days ago | kikifrost | open | No labels | 1 | 0 | 5 days ago |
#3230 New feature: Named Reference Attribution metric — catch citations to the wrong named table, section, or footnote Opened 6 days ago | rajeshl8 | open | No labels | 0 | 0 | 6 days ago |
#3228 MIPROv2's sync execute() calls InstructionProposer.propose() twice, silently doubling LLM cost Opened 6 days ago | Yashwanth-Kumar-Kotla | open | No labels | 0 | 0 | 6 days ago |
#3110 Semantics question: what is a test case's canonical verdict when a judge-based metric is re-run and straddles the threshold? Opened 14 days ago | roy-tong | open | No labels | 7 | 0 | 6 days ago |
#3134 log_hyperparameters documented decorator forms raise TypeError Opened 10 days ago | kikifrost | open | No labels | 2 | 0 | 7 days ago |
#3220 fix(metrics): whitespace-only actual_output/expected_output pass the empty-param guard Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3214 MLLMImage validates the local flag against its URL with a bare assert that vanishes under python -O Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3212 Winogrande benchmark validates n_shots/n_problems with bare asserts and accepts values that crash evaluate() Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3208 HumanEval benchmark validates n/k with a bare assert and accepts invalid values Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3206 Scorer.pass_at_k silently returns wrong scores for degenerate inputs Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3204 CI lint is red on main: `.scripts/release.py` is not black-formatted, blocking every PR Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3200 ArenaTestCase crashes with an opaque IndexError when an arena has fewer than two contestants Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 7 days ago |
#3198 PatternMatchMetric always treats pattern as regex; literal string matching is awkward Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#2903 Return inside finally block suppresses exceptions in async evaluation loop Opened 2 months ago | sahniaditya007 | open | No labels | 2 | 0 | 7 days ago |
#2904 LOCAL_EMBEDDING_API_KEY enum value is defined as a tuple instead of a string Opened 2 months ago | sahniaditya007 | open | No labels | 3 | 0 | 7 days ago |
#2982 Six benchmarks use 'is not' instead of '!=' for length comparison in batch_predict Opened 1 month ago | omsharma0401 | open | No labels | 3 | 0 | 7 days ago |
#3192 fix(metrics): conversational param guard claims 'cannot be empty' but only rejects None Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 7 days ago |
#3190 fix(scorer): replace assert-based argument validation with explicit exceptions Opened 7 days ago | Asthenia0412 | open | No labels | 0 | 0 | 7 days ago |
#3188 fix(scorer): missing optional deps crash with bare NameError instead of a clear error Opened 7 days ago | Asthenia0412 | open | No labels | 1 | 0 | 7 days ago |
#3141 ToolCorrectnessMetric under-scores duplicate tool calls in default (non-exact) mode Opened 9 days ago | Abelo9996 | open | No labels | 2 | 0 | 7 days ago |