ShapeLearn released optimized GGUF quantizations for Qwen 3.8 27B, including versions with embedded MTP draft heads for faster inference. The release includes support for llama.cpp using DFlash2 external drafts for text-only speed optimizations.
HOW THIS AFFECTS YOU
●
builderYou can run highly optimized 27B parameter models on consumer GPUs using speculative decoding via MTP or DFlash2.