Data & Infrastructure

llama.cpp

A C/C++ open-source runtime for running LLM inference on commodity hardware (CPU + Metal / CUDA). Widely used for local Llama, Qwen, DeepSeek, and other open-weights model serving without dedicated inference infrastructure.

Related terms

Framework dimensions

Next steps

Where this fits

llama.cpp is part of the Data & Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.