Repository Issue Activity (beta)

NVIDIA/TensorRT-LLM

Current issue state, recent activity, and per-issue timelines from the indexed issue data.

Open Issues
556
New in 7 Days
31
Closed in 7 Days
106
Average Open Age
131 days
Stale 30+ Days
482
Stale 90+ Days
377
Last 2 Weeks
DateOpenedClosedCommentsEventsOpen Backlog
2026-08-2500000
2026-08-24109133
2026-08-2355286
2026-08-2201256
2026-08-2141123
2026-08-20297001
2026-08-19311917
2026-08-18161184
2026-08-1700011
2026-08-16100110
2026-08-15132055
2026-08-1430122
2026-08-1320056
2026-08-1234381
This Week

Opened: 15

Closed: 105

Comments: 15

Events: 37

Top Labels
triaged (1545)
bug (1136)
stale (711)
feature request (608)
AutoDeploy (477)
question (439)
waiting for feedback (314)
Investigating (259)
Issue Explorer
IssueAuthorStateLabelsCommentsReactionsUpdated

#17964 [Bug]: Qwen3_5ForCausalLM (Qwen3.5-2B) incorrect output with batch_size > 1

Opened 6 days ago
hk0901
open
bug
Customized kernels
Pytorch
9012 hours ago

#18115 [Bug]: V1 MAX_UTILIZATION paused_requests are never processed in the PP executor loop

Opened 23 hours ago
lathuthavarajah
open
Inference runtime
0023 hours ago

#17882 [Bug]: MAX_UTILIZATION resume raises `RequestError: 'modality_type'` for Nemotron Nano 3 Omni

Opened 7 days ago
kota-row
open
bug
LLM API
Multimodal
1024 hours ago

#18104 Withdrawn

Opened 2 days ago
Infatoshi
closed - not_planned
Doc
001 day ago

#18103 Withdrawn

Opened 2 days ago
Infatoshi
closed - not_planned
Triton backend
001 day ago

#18102 Withdrawn

Opened 2 days ago
Infatoshi
closed - not_planned
KV-Cache Management
Speculative Decoding
001 day ago

#18101 Withdrawn

Opened 2 days ago
Infatoshi
closed - not_planned
Memory
001 day ago

#18100 Withdrawn

Opened 2 days ago
Infatoshi
closed - not_planned
KV-Cache Management
001 day ago

#17743 [Bug]: ruff-legacy pre-commit hook reports all baseline violations as regressions on Windows

Opened 9 days ago
edenfunf
open
Windows
002 days ago

#18078 [Bug]: CUDA graphs corrupt generation for google/functiongemma-270m-it (Gemma3 270M) on GB10/sm_121 — argument values collapse; cuda_graph_config: null restores byte-correct output

Opened 3 days ago
yjnhk
closed - not_planned
CUDA Graph
003 days ago

#17665 [Bug]: NIXL cache-transceiver hangs forever if a single completion notification is dropped

Opened 11 days ago
liayan
open
Disaggregated serving
203 days ago

#18085 [RFC] DFlash2 for Qwen3.8 on consumer Blackwell: integration boundary before we propose a PR

Opened 3 days ago
StephenLReed
open
Speculative Decoding
003 days ago

#18084 MODEL_TYPE_TO_TOOL_PARSER maps qwen3_5/qwen3_5_moe to a parser whose format the template does not emit

Opened 3 days ago
StephenLReed
open
LLM API
003 days ago

#18083 Reasoning-parser auto-selection misclassifies reasoning-at-start templates, silently emptying reasoning_content

Opened 3 days ago
StephenLReed
open
LLM API
003 days ago

#14825 [Bug]: kimi 2-6 nvfp4 + SpecD gives low accuracy on GPQA Diamond

Opened 3 months ago
sravan500
closed - completed
bug
Speculative Decoding
303 days ago

#5152 Should remove `self.all_reduce =` from the base class

Opened 1 year ago
yuantailing
closed - completed
triaged
504 days ago

#17024 DeepSeek-V4 + MTP: `AttributeError: '_num_tables'` in `DeepseekV4CacheManager.copy_batch_block_offsets` during warmup

Opened 26 days ago
thorjohnsen
open
No labels
404 days ago

#18039 DeepSeek-V4-Flash NVFP4 baseline aborts in AttentionOp on RTX PRO 6000 Blackwell: Deepseek should be supported by fmha

Opened 4 days ago
rohash123
closed - completed
Pytorch
104 days ago

#10615 [Bug]: Llama 3.2 Vision model fails to build using documented steps

Opened 7 months ago
joy369
closed - completed
bug
Multimodal
804 days ago

#10306 [Usage]: How to use TensorRT LLM to perform MoE NVFP4 | ATTN FP8 inference testing for the Deepseek R1 671B model?

Opened 8 months ago
Alan-D-Chen
closed - completed
question
Inference runtime
104 days ago

#10258 [Bug]: LoRA Support Missing for Encoder-Decoder Models in TensorRT-LLM CPP Implementation

Opened 8 months ago
xiaoxiaoyuwen
closed - completed
bug
Triton backend
Lora/P-tuning
214 days ago

#9523 Support for q4_0 format - google/gemma-3-27b-it-qat-q4_0-unquantized

Opened 9 months ago
albertnanda
closed - completed
question
Low Precision
104 days ago

#8633 [Usage]: TensorRT backend or Python backend

Opened 10 months ago
windtara0619
closed - completed
question
Triton backend
514 days ago

#8572 [Bug]: Build and serving failures with Qwen3-4B-Instruct-2507 model FP8 quantization on H100 GPU with TensorRT-LLM with tensorrt-llm-release:0.19.0 image

Opened 10 months ago
nkg45
closed - completed
bug
Low Precision
Pytorch
104 days ago

#8286 [Usage]: how long the build stage takes for deepseek-r1 models

Opened 11 months ago
ZJLi2013
closed - completed
question
Inference runtime
104 days ago

Rows per page:

1–25 of 3,944