AI Fundamentals

Speculative Decoding

An inference-time optimization that uses a fast draft model to propose tokens which the target model then verifies in parallel, reducing latency without changing output quality. Ships in vLLM, TensorRT-LLM, and most frontier provider stacks.

Related terms

Next steps

Where this fits

Speculative Decoding is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.