AI Fundamentals
Speculative Decoding
An inference-time optimization that uses a fast draft model to propose tokens which the target model then verifies in parallel, reducing latency without changing output quality. Ships in vLLM, TensorRT-LLM, and most frontier provider stacks.
Related terms
Related on this site
Where this fits
Speculative Decoding is part of the AI Fundamentals vocabulary used in the Generative AI Maturity Framework. See the full glossary for the complete set of 149 defined terms, or take the free maturity assessment to see where your organisation stands.