vLLM integrates support for llm-compressor Inkling NVFP4 weights. This allows for the use of 4-bit floating point weights to reduce memory footprint and increase throughput.
HOW THIS AFFECTS YOU
●
builderYou can leverage NVFP4 quantization to run larger models on constrained hardware with reduced VRAM usage.