vllm: Enable GLM-5.2-MXFP4 on the deepseek_v32 path and fix sparse attention correctness
September 22, 2026
vLLM now supports GLM-5.2-MXFP4 via the deepseek_v32 path and includes fixes for sparse attention correctness. These updates enable optimized 4-bit microscaling floating-point inference for the GLM-5.2 architecture while ensuring mathematical accuracy during sparse attention computations.