pytorch/TensorRT

PyTorch/TorchScript/FX compiler for NVIDIA GPUs using TensorRT

View on GitHub ↗Jump to charts ↓Open shareable report →

Data as of . Signed-in members get hourly updates — create a free account.

Summary Information

Updated 1 hour ago
Added to GitGenius on September 23rd, 2026
Created on March 11th, 2020
Open Issues & Pull Requests: 303 (-6)
GitHub issues: Enabled
Number of forks: 414
Total Stargazers: 3,009 (+0)
Total Subscribers: 64 (+0)

Repository Insights (GitGenius)

Median issue/PR response: 2.8 hours
Mean response time: 25.3 days
90th percentile: 14.1 days
Tracked items: 580

How this project is maintained

About 4% of issues opened in the past year have never received a reply. 58% of open issues come from outside the core team, a mix of external reports and the maintainers' own roadmap. 82% of tracked open issues have had no activity in three months, so the open count overstates what is actively being worked. Only 59% of issues opened in the past year have been closed. Three people close 61% of everything that gets resolved.

Charts & Analytics

Fetching additional details & charts...

Issue Activity (beta)

Open issues: 216
New in 7 days: 1
Closed in 7 days: 7
Avg open age: 554 days
Stale 30+ days: 207
Stale 90+ days: 197

Recent activity

Opened in 7 days: 1
Closed in 7 days: 7
Comments in 7 days: 0
Events in 7 days: 6

Top labels

  • bug (331)
  • feature request (86)
  • question (64)
  • story: Operator Coverage & Converters (58)
  • story: Dynamo Frontend & Partitioning (45)
  • story: LLM & Generative AI (35)
  • Story: Runtime & Memory & Serialization (32)
  • Story (30)

Most active issues this week

Sign in to see which issues are moving.
Sign in

Detailed Description

Torch-TensorRT is a compiler that optimizes PyTorch models for inference on NVIDIA GPUs using TensorRT.

The tool addresses the challenge of achieving maximum inference performance on NVIDIA hardware by compiling PyTorch, TorchScript, and FX models through TensorRT's optimization engine. It works by intercepting model execution and applying graph-level optimizations including kernel fusion, precision reduction, and memory layout tuning. The README demonstrates inference latency improvements of up to five times compared to eager execution.

Torch-TensorRT suits teams deploying PyTorch models to NVIDIA GPUs who prioritize inference speed over training flexibility. It works across Linux AMD64, Linux SBSA, Windows, and Jetson platforms, with source compilation supported on Jetson devices. The tool offers two workflows: a torch.compile integration for minimal code changes, and an export-based approach for ahead-of-time optimization and C++ deployment via libtorch. The export workflow enables model serialization for environments without Python dependencies. Support includes optimization techniques like FP8 quantization and is documented with examples for diffusion models, large language models from Hugging Face, and other architectures.

The project maintains active development with nightly builds published alongside stable releases on PyPI, and provides ready-to-run containers through NVIDIA NGC with dependencies and example notebooks included. A formal deprecation policy beginning with version 2.3 communicates API stability expectations to users. The tool is distributed as part of NVIDIA's official PyTorch container ecosystem, indicating integration into the broader NVIDIA platform strategy.