Cerebras delivers Qwen 3.8 27B at 1500 tokens per second
September 3, 2026
Cerebras has added Qwen 3.8 27B to its public endpoints, offering inference speeds of approximately 1500 tokens per second. The model supports 64k context on free tiers and up to 128k on paid tiers using unpruned weights.
HOW THIS AFFECTS YOU
●
builderYou can achieve extremely low-latency inference for mid-sized models using Cerebras's hardware-accelerated endpoints.