Using llama.cpp's DSpark speculative decoding with a 10GB Unsloth drafter, DeepSeek V4 Flash performance increased from 36 to 47 tokens per second on a 7x3090 setup. This achieves near-frontier speeds for local GGUF deployment.
HOW THIS AFFECTS YOU
●
builderYou can significantly improve local inference throughput by utilizing speculative decoding with small drafter models.