vLLM adds speculative decoding support with draft models
September 28, 2026
The vLLM inference runtime now supports speculative decoding using draft models to improve throughput. This optimization allows for faster token generation by using a smaller model to predict tokens before verification by the larger model.
HOW THIS AFFECTS YOU
●
builderYou can reduce inference latency and increase throughput for your LLM deployments using this optimization.
●
researcherThis provides a more efficient way to implement speculative decoding in production-scale inference engines.