Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-25 | 23 | 31 | 75 | 229 | 80 |
| 2026-08-24 | 21 | 19 | 87 | 200 | 90 |
| 2026-08-23 | 24 | 6 | 35 | 125 | 56 |
| 2026-08-22 | 18 | 8 | 37 | 96 | 58 |
| 2026-08-21 | 26 | 26 | 91 | 240 | 94 |
| 2026-08-20 | 34 | 21 | 84 | 283 | 85 |
| 2026-08-19 | 31 | 19 | 106 | 313 | 120 |
| 2026-08-18 | 38 | 15 | 87 | 240 | 110 |
| 2026-08-17 | 29 | 25 | 106 | 282 | 85 |
| 2026-08-16 | 14 | 17 | 51 | 125 | 50 |
| 2026-08-15 | 23 | 3 | 52 | 135 | 54 |
| 2026-08-14 | 24 | 23 | 86 | 194 | 86 |
| 2026-08-13 | 24 | 22 | 84 | 298 | 89 |
| 2026-08-12 | 37 | 24 | 97 | 283 | 99 |
Opened: 177
Closed: 130
Comments: 515
Events: 1,486
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#53748 [Bug]: MLA decode fails on GB10 — shared memory 102400 > 101376; num_stages guard in _decode_grouped_att_m_fwd only fires for BLOCK_DMODEL >= 1024 Opened 9 hours ago | jlacroix82 | open | rocm kimi | 2 | 0 | 7 hours ago |
#53413 [Bug]: GLM-5.2 FP8 on 8×H200 dies with runtime CUDA OOM in sparse_decode_fwd Opened 3 days ago | yuzisun | open | bug glm | 2 | 0 | 7 hours ago |
#49238 [Bug]: Decode instance segfaults on NIXL `loadRemoteMD` after prefill pod restarts in P/D disaggregation Opened 1 month ago | roytman | open | bug | 6 | 0 | 8 hours ago |
#53137 [Bug]: tools/recipes/recipe_json_to_vllm_config.py silently mis-parses the -cc (compilation-config) short alias Opened 5 days ago | wjhrdy | open | rocm | 4 | 0 | 8 hours ago |
#53745 [Bug]: Strict tool calling attaches no structural tag when the reasoning and tool parsers share a parser engine Opened 9 hours ago | csvance | open | bug tool-calling | 0 | 0 | 8 hours ago |
#50587 [Feature]: Kimi K3 Performance Optimization Opened 25 days ago | yewentao256 | open | feature request kimi k3 | 1 | 5 | 8 hours ago |
#50877 [Bug]: DSpark speculative decoding triggers FlashInfer MNNVL allreduce "buffer size insufficient" via draft model's embed_input_ids (TP8, GB200 NVL72) Opened 22 days ago | ilmarkov | open | bug | 2 | 0 | 8 hours ago |
#51581 [Bug][Spec Decode]: DFlash fused-KV projection calls F.linear on a sliced qkv_proj weight — breaks (and can silently corrupt) any weight-quantized drafter Opened 16 days ago | eisbaw | open | quantization | 10 | 1 | 9 hours ago |
#53718 [Bug]: Named shared memory may be closed before peer ranks open it Opened 13 hours ago | rogeroberg | open | bug | 1 | 0 | 9 hours ago |
#53749 [Bug]: enable_prefix_caching=True but zero hits for any shared prefix shorter than one attention block on hybrid checkpoints (2096 / 1920 tokens), with nothing in the logs or metrics naming the minimum Opened 9 hours ago | Schnitzel | open | kimi | 0 | 0 | 9 hours ago |
#51472 [RFC] Multimodal RL inputs for /inference/v1/generate Opened 18 days ago | aoshen02 | open - reopened | RFC multi-modality | 9 | 2 | 9 hours ago |
#53469 [Bug]: structured outputs never enforced for Muse Glimmer with --reasoning-parser muse_glimmer (json_object too; duplicate of #52594) Opened 2 days ago | geekybraindev | open | structured-output | 2 | 0 | 9 hours ago |
#53742 [Feature]: Rust frontend: /v1/embeddings endpoint Opened 9 hours ago | usberkeley | open | feature request rust | 0 | 0 | 9 hours ago |
#51873 [Feature][DSpark]: Enable logprobs with adaptive verification Opened 14 days ago | benchislett | closed - completed | feature request | 1 | 0 | 9 hours ago |
#53737 [Bug]: DeepSeek V3.2 encoder renders a trailing system message inside the assistant think block Opened 10 hours ago | mhuzaifa3 | open | deepseek | 0 | 0 | 10 hours ago |
#44494 [Bug]: Gemma 4 12B is not working Opened 3 months ago | MohamedAliRashad | open | bug | 13 | 2 | 10 hours ago |
#53726 Silent CUDA IMA (exit 0) in hybrid GDN + MTP k=3 + async scheduling on RTX 3090; persists through #50021/#45100/#53613-class fixes Opened 11 hours ago | meowsigma | open | No labels | 0 | 0 | 11 hours ago |
#15081 [Installation]: how to install v0.8.0 Opened 1 year ago | ArlanCooper | closed - completed | installation | 2 | 0 | 11 hours ago |
#53620 [Bug]: test_spec_decode_logprobs never runs on NVIDIA CI, and is ill-defined at top-k ties when it does Opened 1 day ago | jyan-R | open | nvidia | 1 | 0 | 11 hours ago |
#52544 [Bug]: VLLM_USE_PRECOMPILED can fall back across CUDA variants without compatibility validation Opened 9 days ago | jungjiyu | open | bug | 1 | 0 | 11 hours ago |
#53642 [Bug]: Gemma4 parser swallows the rest of the turn when the model emits call:name(...) instead of call:name{...} Opened 1 day ago | chriskosys | closed - completed | bug quantization | 3 | 0 | 11 hours ago |
#35603 [Bug]: vllm: error: unrecognized arguments: --task embedding Opened 6 months ago | dojeese-maker | closed - completed | bug | 13 | 0 | 11 hours ago |
#53665 [Bug]: min_tokens + structured output can empty the token mask and return HTTP 500 Opened 21 hours ago | Yunzez | open | bug structured-output | 3 | 0 | 12 hours ago |
#52803 [ROCm][AMD] Kimi-K3 gfx942 / MI325X Gap and Roadmap Opened 7 days ago | maeehart | open | rocm quantization kimi | 6 | 0 | 12 hours ago |
#53679 [Bug] DSpark/DFlash speculator CUDA graph capture crashes with 'CachingHostAllocator use_count > 0 INTERNAL ASSERT FAILED' after memory profiling Opened 18 hours ago | dsingal0 | open | No labels | 1 | 0 | 13 hours ago |