Llama.cpp Reaches 1.2k t/s Prefill for Qwen3.8 Flash Next
September 12, 2026
A custom llama.cpp branch and HIP runtime optimization achieve 1.2k t/s prefill on Strix Halo hardware, matching the performance of the closed-source Halogen server. This brings open-source prefill speeds to parity with existing proprietary solutions.
HOW THIS AFFECTS YOU
●
builderYou can now achieve high-throughput prefill for Qwen models using open-source tooling on AMD hardware.