llama.cpp v0.6.0 adds MTP speculative decoding for Qwen4Exp
October 5, 2026
The latest llama.cpp release implements Multi-Token Prediction (MTP) speculative decoding specifically for Qwen4Exp models. This update aims to reduce inference latency by predicting multiple tokens in a single forward pass.
HOW THIS AFFECTS YOU
●
builderYou can achieve faster local inference speeds when running Qwen4Exp models on supported hardware.