Maple-Preview achieves 120 tokens per second with ternary 20B MoE on iPhone
August 4, 2026
A ternary 20B Mixture-of-Experts model runs at 120 tokens per second on mobile hardware via Maple-Preview. This implementation leverages low-bit quantization to enable high-throughput inference on an iPhone.
HOW THIS AFFECTS YOU
●
builderYou can explore high-speed local inference for 20B parameter models on mobile devices.
●
founderThis demonstrates a path toward high-performance, edge-deployed AI features without cloud dependency.