Infrastructure

Prefill Caching

The inference optimisation of storing the KV-cache for a prompt prefix so subsequent requests that share that prefix skip recomputation. Distinct from prompt caching (a billing feature): prefill caching is what makes prompt caching possible at the infrastructure level. Governs the cost model of long-context and multi-turn agents.

Last reviewed:

Related terms

Where this fits

Prefill Caching is part of the Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 159 defined terms, or take the free maturity assessment to see where your organisation stands.