Turbo-fieldfare Engine Runs Gemma 4 26B on 2 GB RAM
July 30, 2026
A custom Swift/Metal inference engine enables Gemma 4 26B-A4B-IT to run on Apple Silicon using only 2 GB of RAM. It achieves 5–6 tok/s on an M2 MacBook Air and 31–35 tok/s on M5 hardware, featuring an OpenAI-compatible local server with tool-call support.
HOW THIS AFFECTS YOU
●
builderYou can deploy high-parameter models on low-memory consumer hardware with tool-calling capabilities.
●
researcherThis demonstrates significant efficiency gains in Metal-based quantization and memory management.