Data & Infrastructure

vLLM

A high-throughput open-source LLM inference and serving stack developed at UC Berkeley. Introduced PagedAttention (KV cache paging); the reference implementation for efficient open-weights model serving. Widely used in enterprise self-hosted deployments.

Related terms

Framework dimensions

Next steps

Where this fits

vLLM is part of the Data & Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.