Torch-TensorRT is a compiler that optimizes PyTorch models for inference on NVIDIA GPUs using TensorRT.
The tool addresses the challenge of achieving maximum inference performance on NVIDIA hardware by compiling PyTorch, TorchScript, and FX models through TensorRT's optimization engine. It works by intercepting model execution and applying graph-level optimizations including kernel fusion, precision reduction, and memory layout tuning. The README demonstrates inference latency improvements of up to five times compared to eager execution.
Torch-TensorRT suits teams deploying PyTorch models to NVIDIA GPUs who prioritize inference speed over training flexibility. It works across Linux AMD64, Linux SBSA, Windows, and Jetson platforms, with source compilation supported on Jetson devices. The tool offers two workflows: a torch.compile integration for minimal code changes, and an export-based approach for ahead-of-time optimization and C++ deployment via libtorch. The export workflow enables model serialization for environments without Python dependencies. Support includes optimization techniques like FP8 quantization and is documented with examples for diffusion models, large language models from Hugging Face, and other architectures.
The project maintains active development with nightly builds published alongside stable releases on PyPI, and provides ready-to-run containers through NVIDIA NGC with dependencies and example notebooks included. A formal deprecation policy beginning with version 2.3 communicates API stability expectations to users. The tool is distributed as part of NVIDIA's official PyTorch container ecosystem, indicating integration into the broader NVIDIA platform strategy.