Qwen3.8-Flash-Next Optimized for 1M Context on MLX-serve
September 8, 2026
A specialized MLX-serve implementation enables Qwen3.8-Flash-Next to handle 1M context windows on M5 Max hardware using 8-bit KV cache. It achieves 40 tok/s for prose and 75 tok/s for coding at high context levels using a mixed 8-bit dense and 4-bit expert quantization.
HOW THIS AFFECTS YOU
●
builderYou can now run high-performance, long-context inference locally on Mac hardware with 128GB RAM.