Gemma 4 Achieves 90% Faster Inference via MTP on MLX
June 28, 2026
Gemma 4 inference in Ollama 0.31 on Apple Silicon is up to 90% faster for coding tasks when using multi-token prediction (MTP) powered by MLX. Performance gains were measured using the Aider polyglot benchmark.
HOW THIS AFFECTS YOU
●
builderYou can achieve significantly higher throughput for coding agents on Apple Silicon hardware.