vLLM Fixes Mamba2 Quantized Weight Loading for Tensor Parallelism
September 30, 2026
vLLM updated its inference runtime to correctly load Mamba2 quantized in_proj weights and scales when using tensor parallelism (TP > 1). This fix ensures scaling factors are applied correctly across distributed GPU setups.
HOW THIS AFFECTS YOU
●
builderYou can now reliably deploy quantized Mamba2 models in production using tensor parallelism.
●
researcherThis enables more accurate performance benchmarking of Mamba2 architectures on distributed hardware.