Speculative Decoding Speedups on Consumer Hardware with Device-Agnostic Implementation
July 21, 2026
A new device-agnostic implementation for CUDA, MPS, and CPU demonstrates up to 1.61x wall-clock speedup using K=6 draft tokens. The method maintains distribution equivalence to the target model, verified through two-sample tests on approximately 9,200 tokens.
HOW THIS AFFECTS YOU
●
builderYou can implement faster inference on local hardware like Apple Silicon without sacrificing model output distribution.