Ninfer inference engine delivers 220 tokens per second on Qwen 27B
August 27, 2026
The Ninfer inference server enables throughput of up to 220 tokens per second on a Qwen 27B model using NVFP4 quantization. It supports a 240k context window, FP8 KV cache, and speculative decoding via multi-token prediction.
HOW THIS AFFECTS YOU
●
builderYou can significantly increase inference throughput for large-context models using specialized quantization and speculative decoding.