Qwen3.8-Flash Local Inference Benchmarks on RTX 3090
August 28, 2026
Running Qwen3.8-Flash with IQ4_XS weights and KV-cache quantization yields 160 tok/s prefill and 16 tok/s decode on a single RTX 3090. Using host RAM for experts and n-grams on disk allows the model to fit within 12GB of VRAM.
HOW THIS AFFECTS YOU
●
builderYou can run these models on consumer hardware by utilizing aggressive KV-cache quantization and RAM offloading.