Optimizing Qwen3.8-flash-next on Huawei Ascend 96GB Hardware
October 2, 2026
Using two Huawei Atlas 300I Duo cards, inference performance for Qwen3.8-flash-next improved from 1 token/s to 30 tokens/s for single requests and 61 tokens/s aggregate concurrency. The optimization involved specific modifications to the vLLM and vLLM Ascend software stack.
HOW THIS AFFECTS YOU
●
builderYou can achieve high-throughput inference on relatively inexpensive Huawei Ascend hardware by optimizing the vLLM stack.