Repository Issue Activity (beta)

EleutherAI/lm-evaluation-harness

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
482
New in 7 Days
9
Closed in 7 Days
5
Average Open Age
441 days
Stale 30+ Days
448
Stale 90+ Days
420
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-2400000
2026-08-2300004
2026-08-2220133
2026-08-2133141
2026-08-2012223
2026-08-1920354
2026-08-1810002
2026-08-1700332
2026-08-1620375
2026-08-1510001
2026-08-1400131
2026-08-1300111
2026-08-1200000
2026-08-1100001
This Week

Opened: 9

Closed: 5

Comments: 7

Events: 14

Top Labels
asking questions (74)
bug (67)
feature request (59)
good first issue (29)
help wanted (29)
validation (22)
documentation (6)
opinions wanted (4)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#4038 minerva_math/putnam_axiom/leaderboard-math: thousands-separator comma strip fuses bare digit tuples ("0,1" -> "01")

Opened 2 days ago
feiiiiii5
open
No labels
002 days ago

#4036 minerva_math/putnam_axiom/leaderboard-math: sqrt shorthand normalization corrupts indexed roots \sqrt[3]{x}

Opened 2 days ago
feiiiiii5
open
No labels
002 days ago

#4028 --cache_requests delete fails when the cache directory does not exist

Opened 3 days ago
sdivyanshu90
open
No labels
102 days ago

#3046 Cache fails when repeats are greater than 1

Opened 1 year ago
eyuansu62
open
bug
112 days ago

#3749 Feature request: canonical reproducibility-bundle export in scripts/

Opened 4 months ago
topeuph-ai
open
No labels
1202 days ago

#4031 minerva_math: REMOVED_EXPRESSIONS entry "ft" corrupts \left (substring replace inside control sequences)

Opened 3 days ago
rcfa
open
No labels
003 days ago

#4030 minerva_math: identical answers scored wrong when sympy cannot parse them (tuples, intervals, matrices)

Opened 3 days ago
rcfa
open
No labels
003 days ago

#4002 Scalar `seed` in a YAML config crashes with TypeError; `seed: 0` silently disables seeding

Opened 9 days ago
winklemad
closed - completed
No labels
003 days ago

#4011 `_parse_dict_args` is a no-op under postponed annotations, and `DICT_KEYS` is unused

Opened 6 days ago
winklemad
closed - completed
No labels
003 days ago

#3962 [Proposal] OpenEval Import/Export Support

Opened 26 days ago
adhabnr-ux
closed - not_planned
No labels
203 days ago

#3558 think_end_token support needed for local-chat-completions mode

Opened 7 months ago
kediaharshit9
closed - completed
No labels
124 days ago

#3460 max_model_length for gemma3-12b is set to 2048 (_DEFAULT_MAX_LENGTH) with the hf processor

Opened 9 months ago
bodasadallah
closed - completed
No labels
004 days ago

#4022 Feature request: EvalPort import/export for per-document samples and results

Opened 4 days ago
adhabnr-ux
open
No labels
004 days ago

#4019 Proposal: Wilson CIs + content-hash integrity field on results

Opened 5 days ago
CSOAI-ORG
open
No labels
105 days ago

#3017 Support for using a remote /tokenize API endpoint as the tokenizer

Opened 1 year ago
furkancoskun
open
feature request
245 days ago

#4017 Proposal: Wilson CI bounds, MDE, and optional BH-FDR flags on results output

Opened 5 days ago
CSOAI-ORG
open
No labels
005 days ago

#4007 Generative-task scorers cannot distinguish an unparseable response from a wrong answer (census: 2,936/4,524 tasks affected)

Opened 8 days ago
Rajveer-code
open
No labels
207 days ago

#4005 --use_cache crashes before inference for multimodal image requests

Opened 8 days ago
sdivyanshu90
open
No labels
008 days ago

#3080 Fix Metric Calculation for Repeats

Opened 1 year ago
baberabb
open
No labels
608 days ago

#3493 Standardizing Task Formats

Opened 7 months ago
baberabb
open
No labels
109 days ago

#3881 Request cache key ignores generation_kwargs -> silently reuses cached instances across different sampling parameters

Opened 2 months ago
AUTHENSOR
open
No labels
1010 days ago

#3498 what's the proper config for aime24 evalated on Qwen3?

Opened 7 months ago
Lynnzake
open
No labels
2211 days ago

#3989 Fix broken Discord link in contributing guide

Opened 14 days ago
sdivyanshu90
open
No labels
0014 days ago

#3930 [Bug] afrixnli prompt_1: doc_to_text uses .format() braces, so premise and hypothesis are never substituted (35 tasks)

Opened 1 month ago
AmosBunde
closed - completed
No labels
2014 days ago

#3859 SCROLLS tasks fail to load with recent datasets (dataset scripts no longer supported)

Opened 2 months ago
bongho
closed - completed
No labels
2015 days ago

Rows per page:

1–25 of 922