NVIDIA's NVFP4 quantization reduces the 125B MoE Qwen3.8-Flash-Next model size by 63% while utilizing hybrid attention. The weights are now available on Hugging Face with minimal accuracy degradation compared to the original precision.
HOW THIS AFFECTS YOU
●
builderYou can deploy larger MoE models with significantly lower memory footprints on existing hardware.
●
researcherThis provides a new benchmark for FP4 quantization effectiveness in large mixture-of-experts architectures.