Data & Infrastructure

Prompt Caching

A provider-side optimization that reuses the KV cache for repeated prompt prefixes across requests, typically reducing input token cost by ~10× for reused content. Anthropic, OpenAI, and Google offer it. Order-of-magnitude cost driver for RAG and agent workloads.

Related terms

Framework dimensions

Next steps

Where this fits

Prompt Caching is part of the Data & Infrastructure vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.