Current issue state, recent activity, and per-issue timelines from the indexed issue data.
| Date | Opened | Closed | Comments | Events | Open Backlog |
|---|---|---|---|---|---|
| 2026-08-25 | 0 | 0 | 0 | 0 | 0 |
| 2026-08-24 | 1 | 0 | 9 | 13 | 3 |
| 2026-08-23 | 5 | 5 | 2 | 8 | 6 |
| 2026-08-22 | 0 | 1 | 2 | 5 | 6 |
| 2026-08-21 | 4 | 1 | 1 | 2 | 3 |
| 2026-08-20 | 2 | 97 | 0 | 0 | 1 |
| 2026-08-19 | 3 | 1 | 1 | 9 | 17 |
| 2026-08-18 | 16 | 1 | 1 | 8 | 4 |
| 2026-08-17 | 0 | 0 | 0 | 1 | 1 |
| 2026-08-16 | 1 | 0 | 0 | 1 | 10 |
| 2026-08-15 | 13 | 2 | 0 | 5 | 5 |
| 2026-08-14 | 3 | 0 | 1 | 2 | 2 |
| 2026-08-13 | 2 | 0 | 0 | 5 | 6 |
| 2026-08-12 | 3 | 4 | 3 | 8 | 1 |
Opened: 15
Closed: 105
Comments: 15
Events: 37
| Issue | Author | State | Labels | Comments | Reactions | Updated |
|---|---|---|---|---|---|---|
#17964 [Bug]: Qwen3_5ForCausalLM (Qwen3.5-2B) incorrect output with batch_size > 1 Opened 6 days ago | hk0901 | open | bug Customized kernels Pytorch | 9 | 0 | 12 hours ago |
#18115 [Bug]: V1 MAX_UTILIZATION paused_requests are never processed in the PP executor loop Opened 23 hours ago | lathuthavarajah | open | Inference runtime | 0 | 0 | 23 hours ago |
#17882 [Bug]: MAX_UTILIZATION resume raises `RequestError: 'modality_type'` for Nemotron Nano 3 Omni Opened 7 days ago | kota-row | open | bug LLM API Multimodal | 1 | 0 | 24 hours ago |
#18104 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | Doc | 0 | 0 | 1 day ago |
#18103 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | Triton backend | 0 | 0 | 1 day ago |
#18102 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | KV-Cache Management Speculative Decoding | 0 | 0 | 1 day ago |
#18101 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | Memory | 0 | 0 | 1 day ago |
#18100 Withdrawn Opened 2 days ago | Infatoshi | closed - not_planned | KV-Cache Management | 0 | 0 | 1 day ago |
#17743 [Bug]: ruff-legacy pre-commit hook reports all baseline violations as regressions on Windows Opened 9 days ago | edenfunf | open | Windows | 0 | 0 | 2 days ago |
#18078 [Bug]: CUDA graphs corrupt generation for google/functiongemma-270m-it (Gemma3 270M) on GB10/sm_121 — argument values collapse; cuda_graph_config: null restores byte-correct output Opened 3 days ago | yjnhk | closed - not_planned | CUDA Graph | 0 | 0 | 3 days ago |
#17665 [Bug]: NIXL cache-transceiver hangs forever if a single completion notification is dropped Opened 11 days ago | liayan | open | Disaggregated serving | 2 | 0 | 3 days ago |
#18085 [RFC] DFlash2 for Qwen3.8 on consumer Blackwell: integration boundary before we propose a PR Opened 3 days ago | StephenLReed | open | Speculative Decoding | 0 | 0 | 3 days ago |
#18084 MODEL_TYPE_TO_TOOL_PARSER maps qwen3_5/qwen3_5_moe to a parser whose format the template does not emit Opened 3 days ago | StephenLReed | open | LLM API | 0 | 0 | 3 days ago |
#18083 Reasoning-parser auto-selection misclassifies reasoning-at-start templates, silently emptying reasoning_content Opened 3 days ago | StephenLReed | open | LLM API | 0 | 0 | 3 days ago |
#14825 [Bug]: kimi 2-6 nvfp4 + SpecD gives low accuracy on GPQA Diamond Opened 3 months ago | sravan500 | closed - completed | bug Speculative Decoding | 3 | 0 | 3 days ago |
#5152 Should remove `self.all_reduce =` from the base class Opened 1 year ago | yuantailing | closed - completed | triaged | 5 | 0 | 4 days ago |
#17024 DeepSeek-V4 + MTP: `AttributeError: '_num_tables'` in `DeepseekV4CacheManager.copy_batch_block_offsets` during warmup Opened 26 days ago | thorjohnsen | open | No labels | 4 | 0 | 4 days ago |
#18039 DeepSeek-V4-Flash NVFP4 baseline aborts in AttentionOp on RTX PRO 6000 Blackwell: Deepseek should be supported by fmha Opened 4 days ago | rohash123 | closed - completed | Pytorch | 1 | 0 | 4 days ago |
#10615 [Bug]: Llama 3.2 Vision model fails to build using documented steps Opened 7 months ago | joy369 | closed - completed | bug Multimodal | 8 | 0 | 4 days ago |
#10306 [Usage]: How to use TensorRT LLM to perform MoE NVFP4 | ATTN FP8 inference testing for the Deepseek R1 671B model? Opened 8 months ago | Alan-D-Chen | closed - completed | question Inference runtime | 1 | 0 | 4 days ago |
#10258 [Bug]: LoRA Support Missing for Encoder-Decoder Models in TensorRT-LLM CPP Implementation Opened 8 months ago | xiaoxiaoyuwen | closed - completed | bug Triton backend Lora/P-tuning | 2 | 1 | 4 days ago |
#9523 Support for q4_0 format - google/gemma-3-27b-it-qat-q4_0-unquantized Opened 9 months ago | albertnanda | closed - completed | question Low Precision | 1 | 0 | 4 days ago |
#8633 [Usage]: TensorRT backend or Python backend Opened 10 months ago | windtara0619 | closed - completed | question Triton backend | 5 | 1 | 4 days ago |
#8572 [Bug]: Build and serving failures with Qwen3-4B-Instruct-2507 model FP8 quantization on H100 GPU with TensorRT-LLM with tensorrt-llm-release:0.19.0 image Opened 10 months ago | nkg45 | closed - completed | bug Low Precision Pytorch | 1 | 0 | 4 days ago |
#8286 [Usage]: how long the build stage takes for deepseek-r1 models Opened 11 months ago | ZJLi2013 | closed - completed | question Inference runtime | 1 | 0 | 4 days ago |