Repository Issue Activity (beta)

open-compass/vlmevalkit

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
220
New in 7 Days
2
Closed in 7 Days
2
Average Open Age
431 days
Stale 30+ Days
215
Stale 90+ Days
202
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-09-2001010
2026-09-1900000
2026-09-1801020
2026-09-1700111
2026-09-1620112
2026-09-150000220
2026-09-1400000
2026-09-1300000
2026-09-1200000
2026-09-1100000
2026-09-1000000
2026-09-0910000
2026-09-0810000
2026-09-0700000
This Week

Opened: 2

Closed: 2

Comments: 2

Events: 5

Top Labels
Awaiting Confirm (17)
Feature Request (16)
WIP (3)
BUG (2)
Further Info (1)
Pending (1)
Won't Fix (1)
documentation (1)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#1679 MathVista post_check: non-numeric answers can never be scored correct (str(res) typo, also in qspatial)

Opened 12 days ago
rrrxxx0510
closed - completed
No labels
002 days ago

#1690 olmOCRBench: `evaluate()` returns `None`, so the score never reaches `status.json`

Opened 5 days ago
jaeminSon
open
No labels
304 days ago

#1689 `dump()` silently drops predictions that begin with a URL and exceed 2079 characters

Opened 5 days ago
jaeminSon
open
No labels
005 days ago

#1661 OpenAI-compat: GPT4V api_base is the full /v1/chat/completions URL, not /v1

Opened 24 days ago
cursor[bot]
open
No labels
208 days ago

#1675 DOCLING / IDEFICS2 / Mantis / XGenMM fail to import on transformers 5.x: AutoModelForVision2Seq was removed

Opened 13 days ago
rrrxxx0510
open
No labels
0013 days ago

#1665 Structured extra_records output is discarded for the entire dataset if any single sample fails

Opened 20 days ago
williamwuyantao
open
No labels
1013 days ago

#1652 Interop idea: a VLMEvalKit-openeval-adapter for portable eval results

Opened 28 days ago
adhabnr-ux
open
No labels
0028 days ago

#1590 [Bug] Thyme: sandbox image fed in assistant turn + KV-cache reused across re-encoded turns corrupts multi-round vision context

Opened 3 months ago
Irisicy4
open
No labels
101 month ago

#323 [Help Wanted] Supporting the `chat_inner` API for existing VLMs.

Opened 2 years ago
kennymckormick
open
Feature Request
301 month ago

#1632 Add Pluto-AI-Labs/Apollo-VL-Edge-3B to Leaderboard

Opened 1 month ago
SiddharthFX
open
No labels
001 month ago

#1614 有些基准没有放数据url

Opened 2 months ago
yangyangyang127
open
No labels
001 month ago

#1621 Qwen2.5-VL-7B模型,对评测集LogicVista,从推理结果中提取答案错误,导致误判

Opened 2 months ago
suddenlee
open
No labels
002 months ago

#1617 MMMU Open-ended问题无法正确评测

Opened 2 months ago
identzz-ai
open - reopened
No labels
002 months ago

#1606 VLM2Bench counting metric silently degenerates to exact match (image_seq_len always defaults to 2)

Opened 2 months ago
veerendrav
closed - completed
No labels
002 months ago

#1615 OpenVLM: one-decimal Overall makes the board un-checkable, and hides that 8 boards' #1 leads are under 1.2 sigma

Opened 2 months ago
ipezygj
open
No labels
102 months ago

#1549 RFC: Ascend NPU device support

Opened 4 months ago
pjgao
closed - completed
No labels
122 months ago

#1600 LiveMMBench空链接

Opened 2 months ago
yangyangyang127
open
No labels
002 months ago

#1456 【OmniDocBench】Recommend updating the benchmark version

Opened 7 months ago
TMYsecret
open
No labels
503 months ago

#1598 请问flyai-vl模型的评测什么时候可以支持?

Opened 3 months ago
zhangxuewei-xirui
open
No labels
003 months ago

#1589 [Bug] run.py: LOCAL_RANK defaults to 1 instead of 0

Opened 3 months ago
Irisicy4
open
No labels
103 months ago

#1563 LongVideoBench `use_subtitle` logic seems inverted

Opened 4 months ago
zyuhan1999
closed - completed
No labels
203 months ago

#1582 Adding [BenchCAD]: would you accept execution-based image→CadQuery-code tasks (heavier deps)?

Opened 3 months ago
HaozheZhang6
open
No labels
003 months ago

#1580 [Bug] CharXiv_descriptive_val DATASET_MD5 appears to be incorrect after #1555

Opened 3 months ago
kail8
open
No labels
003 months ago

#1567 MMBench_v11 md5 failed

Opened 4 months ago
Uradouby
closed - completed
No labels
103 months ago

#1516 About MMBench Evaluation

Opened 5 months ago
cfmaXT
open
No labels
333 months ago

Rows per page:

1–25 of 484