[r/LocalLLaMA]score: 0.14
Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak
August 16, 2026
An optimized vLLM implementation for Qwen2.5-27B achieves 82 tps single-request and 417 tps sustained throughput on a single RTX 3090. By combining W4A16 weights with FP8 KV cache and INT8 quantization for lm_head and embed_tokens, the engine supports up to 200k context within 14.2GB of VRAM.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy