Repository Issue Activity (beta)

vllm-project/vllm

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
2,215
New in 7 Days
215
Closed in 7 Days
145
Average Open Age
55 days
Stale 30+ Days
912
Stale 90+ Days
66
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-2523317522980
2026-08-2421198720090
2026-08-232463512556
2026-08-22188379658
2026-08-2126269124094
2026-08-2034218428385
2026-08-193119106313120
2026-08-18381587240110
2026-08-17292510628285
2026-08-1614175112550
2026-08-152335213554
2026-08-1424238619486
2026-08-1324228429889
2026-08-1237249728399
This Week

Opened: 177

Closed: 130

Comments: 515

Events: 1,486

Top Labels
bug (7810)
stale (6970)
feature request (2072)
usage (1697)
RFC (620)
performance (516)
rocm (473)
installation (471)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#53748 [Bug]: MLA decode fails on GB10 — shared memory 102400 > 101376; num_stages guard in _decode_grouped_att_m_fwd only fires for BLOCK_DMODEL >= 1024

Opened 9 hours ago
jlacroix82
open
rocm
kimi
207 hours ago

#53413 [Bug]: GLM-5.2 FP8 on 8×H200 dies with runtime CUDA OOM in sparse_decode_fwd

Opened 3 days ago
yuzisun
open
bug
glm
207 hours ago

#49238 [Bug]: Decode instance segfaults on NIXL `loadRemoteMD` after prefill pod restarts in P/D disaggregation

Opened 1 month ago
roytman
open
bug
608 hours ago

#53137 [Bug]: tools/recipes/recipe_json_to_vllm_config.py silently mis-parses the -cc (compilation-config) short alias

Opened 5 days ago
wjhrdy
open
rocm
408 hours ago

#53745 [Bug]: Strict tool calling attaches no structural tag when the reasoning and tool parsers share a parser engine

Opened 9 hours ago
csvance
open
bug
tool-calling
008 hours ago

#50587 [Feature]: Kimi K3 Performance Optimization

Opened 25 days ago
yewentao256
open
feature request
kimi
k3
158 hours ago

#50877 [Bug]: DSpark speculative decoding triggers FlashInfer MNNVL allreduce "buffer size insufficient" via draft model's embed_input_ids (TP8, GB200 NVL72)

Opened 22 days ago
ilmarkov
open
bug
208 hours ago

#51581 [Bug][Spec Decode]: DFlash fused-KV projection calls F.linear on a sliced qkv_proj weight — breaks (and can silently corrupt) any weight-quantized drafter

Opened 16 days ago
eisbaw
open
quantization
1019 hours ago

#53718 [Bug]: Named shared memory may be closed before peer ranks open it

Opened 13 hours ago
rogeroberg
open
bug
109 hours ago

#53749 [Bug]: enable_prefix_caching=True but zero hits for any shared prefix shorter than one attention block on hybrid checkpoints (2096 / 1920 tokens), with nothing in the logs or metrics naming the minimum

Opened 9 hours ago
Schnitzel
open
kimi
009 hours ago

#51472 [RFC] Multimodal RL inputs for /inference/v1/generate

Opened 18 days ago
aoshen02
open - reopened
RFC
multi-modality
929 hours ago

#53469 [Bug]: structured outputs never enforced for Muse Glimmer with --reasoning-parser muse_glimmer (json_object too; duplicate of #52594)

Opened 2 days ago
geekybraindev
open
structured-output
209 hours ago

#53742 [Feature]: Rust frontend: /v1/embeddings endpoint

Opened 9 hours ago
usberkeley
open
feature request
rust
009 hours ago

#51873 [Feature][DSpark]: Enable logprobs with adaptive verification

Opened 14 days ago
benchislett
closed - completed
feature request
109 hours ago

#53737 [Bug]: DeepSeek V3.2 encoder renders a trailing system message inside the assistant think block

Opened 10 hours ago
mhuzaifa3
open
deepseek
0010 hours ago

#44494 [Bug]: Gemma 4 12B is not working

Opened 3 months ago
MohamedAliRashad
open
bug
13210 hours ago

#53726 Silent CUDA IMA (exit 0) in hybrid GDN + MTP k=3 + async scheduling on RTX 3090; persists through #50021/#45100/#53613-class fixes

Opened 11 hours ago
meowsigma
open
No labels
0011 hours ago

#15081 [Installation]: how to install v0.8.0

Opened 1 year ago
ArlanCooper
closed - completed
installation
2011 hours ago

#53620 [Bug]: test_spec_decode_logprobs never runs on NVIDIA CI, and is ill-defined at top-k ties when it does

Opened 1 day ago
jyan-R
open
nvidia
1011 hours ago

#52544 [Bug]: VLLM_USE_PRECOMPILED can fall back across CUDA variants without compatibility validation

Opened 9 days ago
jungjiyu
open
bug
1011 hours ago

#53642 [Bug]: Gemma4 parser swallows the rest of the turn when the model emits call:name(...) instead of call:name{...}

Opened 1 day ago
chriskosys
closed - completed
bug
quantization
3011 hours ago

#35603 [Bug]: vllm: error: unrecognized arguments: --task embedding

Opened 6 months ago
dojeese-maker
closed - completed
bug
13011 hours ago

#53665 [Bug]: min_tokens + structured output can empty the token mask and return HTTP 500

Opened 21 hours ago
Yunzez
open
bug
structured-output
3012 hours ago

#52803 [ROCm][AMD] Kimi-K3 gfx942 / MI325X Gap and Roadmap

Opened 7 days ago
maeehart
open
rocm
quantization
kimi
6012 hours ago

#53679 [Bug] DSpark/DFlash speculator CUDA graph capture crashes with 'CachingHostAllocator use_count > 0 INTERNAL ASSERT FAILED' after memory profiling

Opened 18 hours ago
dsingal0
open
No labels
1013 hours ago

Rows per page:

1–25 of 17,059