Skip to main content
llm.info

TensorRT-LLM

Open Source
Library

Optimized LLM inference on NVIDIA GPUs

About

TensorRT-LLM is NVIDIA's open library for compiling and serving LLMs with maximum throughput and low latency on NVIDIA GPUs, using kernel fusion, quantization, in-flight batching, and multi-GPU parallelism.

Compatibility

Supported Languages

python
c++

Details

Category
Library
License
Apache-2.0