vLLM Extends Device-Side MM Normalization to GLM4V and GLM5Next
September 30, 2026
vLLM has added device-side multimodal (MM) normalization support for the GLM4V and GLM5Next model families. This optimization moves normalization computations to the GPU to reduce CPU bottlenecks.
HOW THIS AFFECTS YOU
●
builderYou can achieve lower latency when serving GLM-based multimodal models by offloading normalization to the device.