W8A8 Kernels Enable 1.4x Speedup on Apple M5 Silicon
July 23, 2026
Custom W8A8 kernels for Apple M5 hardware achieve 3,029 tps for Gemma4 prefill tasks, a 1.4x improvement over the 2,193 tps baseline. Current inference backends like MLX and Llama.cpp lack support for the M5's native INT8 activation capabilities.
HOW THIS AFFECTS YOU
●
builderYou can achieve significantly higher prefill throughput on M5 hardware by implementing INT8 activation kernels.