Bypassing Apple M3 Neural Engine DRAM Bottlenecks to Double Token Throughput
September 9, 2026
An RTL erratum in the M3 Neural Engine throttles DRAM weight streaming to 17-19 GB/s when weight sizes are integer multiples of 1 MiB. By avoiding the problematic speculative prefetch ring in the kernel DMA engine, Llama 3.2 1B throughput increases from 10.0 to 24.3 tokens/s.
HOW THIS AFFECTS YOU
●
builderYou can achieve over 2x performance gains on M3 hardware by adjusting kernel DMA prefetch paths or model weight dimensions.