Strata inference engine achieves 50t/s on Qwen3.8-Flash-Next
September 30, 2026
The Strata inference engine demonstrates 50t/s token generation and 1500t/s prompt processing on a 12GB VRAM laptop using Qwen3.8-Flash-Next GGUF. Currently, the engine is optimized for NVIDIA hardware with experimental AMD support.
HOW THIS AFFECTS YOU
●
builderYou can achieve significantly higher inference speeds on consumer hardware compared to standard llama.cpp forks.