Data & Infrastructure

TensorRT-LLM

NVIDIA's inference stack for LLMs, optimized for NVIDIA GPUs. Includes quantization, speculative decoding, and multi-GPU tensor parallelism. Fastest option for open-weights inference on NVIDIA hardware; less flexible on other accelerators.

Related terms

Framework dimensions

Next steps

Where this fits

TensorRT-LLM is part of the Data & Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.