Strix Halo Optimized Inference Achieves 1300tok/s Prefill for Qwen 3.8 Flash Next
September 13, 2026
Custom llama.cpp forks and halogen-flash-server on Strix Halo hardware achieve 52tok/s decode and 1300tok/s prefill for Qwen 3.8 Flash Next. These optimizations leverage Engram architecture improvements to maximize performance on constrained local compute setups.
HOW THIS AFFECTS YOU
●
builderYou can achieve significantly higher throughput on local hardware by utilizing these specific inference engine forks.
●
researcherThe Engram architecture in Qwen 3.8 Flash Next demonstrates high efficiency for small-scale deployments.