Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-09-20 | 0 | 1 | 0 | 1 | 0 |
| 2026-09-19 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-18 | 0 | 1 | 0 | 2 | 0 |
| 2026-09-17 | 0 | 0 | 1 | 1 | 1 |
| 2026-09-16 | 2 | 0 | 1 | 1 | 2 |
| 2026-09-15 | 0 | 0 | 0 | 0 | 220 |
| 2026-09-14 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-13 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-12 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-11 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-10 | 0 | 0 | 0 | 0 | 0 |
| 2026-09-09 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-08 | 1 | 0 | 0 | 0 | 0 |
| 2026-09-07 | 0 | 0 | 0 | 0 | 0 |
Opened: 2
Closed: 2
Comments: 2
Events: 5
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#1679 MathVista post_check: non-numeric answers can never be scored correct (str(res) typo, also in qspatial) Opened 12 days ago | rrrxxx0510 | closed - completed | No labels | 0 | 0 | 2 days ago |
#1690 olmOCRBench: `evaluate()` returns `None`, so the score never reaches `status.json` Opened 5 days ago | jaeminSon | open | No labels | 3 | 0 | 4 days ago |
#1689 `dump()` silently drops predictions that begin with a URL and exceed 2079 characters Opened 5 days ago | jaeminSon | open | No labels | 0 | 0 | 5 days ago |
#1661 OpenAI-compat: GPT4V api_base is the full /v1/chat/completions URL, not /v1 Opened 24 days ago | cursor[bot] | open | No labels | 2 | 0 | 8 days ago |
#1675 DOCLING / IDEFICS2 / Mantis / XGenMM fail to import on transformers 5.x: AutoModelForVision2Seq was removed Opened 13 days ago | rrrxxx0510 | open | No labels | 0 | 0 | 13 days ago |
#1665 Structured extra_records output is discarded for the entire dataset if any single sample fails Opened 20 days ago | williamwuyantao | open | No labels | 1 | 0 | 13 days ago |
#1652 Interop idea: a VLMEvalKit-openeval-adapter for portable eval results Opened 28 days ago | adhabnr-ux | open | No labels | 0 | 0 | 28 days ago |
#1590 [Bug] Thyme: sandbox image fed in assistant turn + KV-cache reused across re-encoded turns corrupts multi-round vision context Opened 3 months ago | Irisicy4 | open | No labels | 1 | 0 | 1 month ago |
#323 [Help Wanted] Supporting the `chat_inner` API for existing VLMs. Opened 2 years ago | kennymckormick | open | Feature Request | 3 | 0 | 1 month ago |
#1632 Add Pluto-AI-Labs/Apollo-VL-Edge-3B to Leaderboard Opened 1 month ago | SiddharthFX | open | No labels | 0 | 0 | 1 month ago |
#1614 有些基准没有放数据url Opened 2 months ago | yangyangyang127 | open | No labels | 0 | 0 | 1 month ago |
#1621 Qwen2.5-VL-7B模型,对评测集LogicVista,从推理结果中提取答案错误,导致误判 Opened 2 months ago | suddenlee | open | No labels | 0 | 0 | 2 months ago |
#1617 MMMU Open-ended问题无法正确评测 Opened 2 months ago | identzz-ai | open - reopened | No labels | 0 | 0 | 2 months ago |
#1606 VLM2Bench counting metric silently degenerates to exact match (image_seq_len always defaults to 2) Opened 2 months ago | veerendrav | closed - completed | No labels | 0 | 0 | 2 months ago |
#1615 OpenVLM: one-decimal Overall makes the board un-checkable, and hides that 8 boards' #1 leads are under 1.2 sigma Opened 2 months ago | ipezygj | open | No labels | 1 | 0 | 2 months ago |
#1549 RFC: Ascend NPU device support Opened 4 months ago | pjgao | closed - completed | No labels | 1 | 2 | 2 months ago |
#1600 LiveMMBench空链接 Opened 2 months ago | yangyangyang127 | open | No labels | 0 | 0 | 2 months ago |
#1456 【OmniDocBench】Recommend updating the benchmark version Opened 7 months ago | TMYsecret | open | No labels | 5 | 0 | 3 months ago |
#1598 请问flyai-vl模型的评测什么时候可以支持? Opened 3 months ago | zhangxuewei-xirui | open | No labels | 0 | 0 | 3 months ago |
#1589 [Bug] run.py: LOCAL_RANK defaults to 1 instead of 0 Opened 3 months ago | Irisicy4 | open | No labels | 1 | 0 | 3 months ago |
#1563 LongVideoBench `use_subtitle` logic seems inverted Opened 4 months ago | zyuhan1999 | closed - completed | No labels | 2 | 0 | 3 months ago |
#1582 Adding [BenchCAD]: would you accept execution-based image→CadQuery-code tasks (heavier deps)? Opened 3 months ago | HaozheZhang6 | open | No labels | 0 | 0 | 3 months ago |
#1580 [Bug] CharXiv_descriptive_val DATASET_MD5 appears to be incorrect after #1555 Opened 3 months ago | kail8 | open | No labels | 0 | 0 | 3 months ago |
#1567 MMBench_v11 md5 failed Opened 4 months ago | Uradouby | closed - completed | No labels | 1 | 0 | 3 months ago |
#1516 About MMBench Evaluation Opened 5 months ago | cfmaXT | open | No labels | 3 | 3 | 3 months ago |