TensorRT-LLM
Open Source
Library
Optimized LLM inference on NVIDIA GPUs
About
TensorRT-LLM is NVIDIA's open library for compiling and serving LLMs with maximum throughput and low latency on NVIDIA GPUs, using kernel fusion, quantization, in-flight batching, and multi-GPU parallelism.
Compatibility
Supported Languages
python
c++
Details
- Category
- Library
- License
- Apache-2.0