Repository Issue Activity (beta)

openai/simple-evals

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
33
New in 7 Days
0
Closed in 7 Days
0
Average Open Age
431 days
Stale 30+ Days
31
Stale 90+ Days
28
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-1500000
2026-09-14000033
2026-09-1300000
2026-09-1200000
2026-09-1100000
2026-09-1000000
2026-09-0900000
2026-09-0800000
2026-09-0700000
2026-09-0600000
2026-09-0500000
2026-09-0400000
2026-09-0300000
2026-09-0200000
This Week

Opened: 0

Closed: 0

Comments: 0

Events: 0

Top Labels

No label distribution is available yet.

Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#120 OpenAI-compat: ChatCompletionSampler plus OPENAI_BASE_URL /v1

Opened 18 days ago
cursor[bot]
open
No labels
0018 days ago

#117 Optional EvalPort interop for SingleEvalResult / EvalResult

Opened 24 days ago
adhabnr-ux
open
No labels
0024 days ago

#116 Reproducibility: provenance bundle + paired statistics for fixed-item evals

Opened 1 month ago
KeilerHirsch
open
No labels
001 month ago

#115 MATH grading takes the first "Answer:" match (not the final answer), and check_equality interpolates candidate text raw into the grader prompt

Opened 1 month ago
AUTHENSOR
open
No labels
001 month ago

#83 Medical category of questions in the dataset

Opened 1 year ago
sid-artpark
open
No labels
133 months ago

#109 Standard deviations for healthbench are wrong

Opened 4 months ago
za3k
open
No labels
004 months ago

#103 Question: where would a long-horizon “tension crash test” benchmark best live in your eval ecosystem?

Opened 7 months ago
onestardao
closed - completed
No labels
106 months ago

#105 Batching for higher throughput?

Opened 6 months ago
szhou0202
open
No labels
006 months ago

#102 Question about healthbench mandarin version

Opened 8 months ago
sbrz13
closed - completed
No labels
008 months ago

#100 Regarding the environmental data issue of the web search results of BrowseComp?

Opened 9 months ago
whfeLingYu
open
No labels
029 months ago

#99 Answer of a question in browsecomp may be wrong

Opened 10 months ago
flibbertigibbet-Y
open
No labels
0010 months ago

#98 What dataset did math use?

Opened 11 months ago
yongho-chang
open
No labels
0011 months ago

#97 What to do now that simple-evals won't be updated?

Opened 11 months ago
william-max-byte
open
No labels
2011 months ago

#96 healthbench_eval can not reproduced‌ 0.67 on the gpt-5

Opened 1 year ago
hitwangshuai
open
No labels
101 year ago

#7 Run benchmarks also for GPT-3.5 versions and Claude Sonnet and Haiku

Opened 2 years ago
zurferr
closed - completed
No labels
101 year ago

#85 Cost of a Single Evaluation Using GPT-4.1 in HealthBench

Opened 1 year ago
Rorschaaaach
open
No labels
101 year ago

#90 Possible Typo in DROP Benchmark Accuracy

Opened 1 year ago
IkerJansa44
open
No labels
011 year ago

#89 GPQA Accuracy Mismatch on Completions vs Responses API (GPT-4o August 2024)

Opened 1 year ago
ziwenseal
open
No labels
001 year ago

#88 Environment setup file for Healthbench

Opened 1 year ago
drixs2050
open
No labels
001 year ago

#28 How do we run this code?

Opened 2 years ago
atahanuz
open
No labels
311 year ago

#80 HealthBench asking for azure credentials

Opened 1 year ago
shaafsalman
closed - completed
No labels
301 year ago

#79 Clearly wrong / tampered prompt_id with "milf"

Opened 1 year ago
junghoon-son
open
No labels
161 year ago

#78 HealthBench performance data json

Opened 1 year ago
yzlnew
open
No labels
111 year ago

#77 Incorrect scores in HealthBench?

Opened 1 year ago
davidgilbertson
open
No labels
401 year ago

#76 Access to HealthBench Dataset?

Opened 1 year ago
alt164
open
No labels
231 year ago

Rows per page:

1–25 of 38