Mference Engine Enables 284B DeepSeek-V4-Flash on 5.3GB Memory
August 2, 2026
The Mference engine uses a streaming approach for MoE models, keeping the shared core and KV cache resident while streaming experts from SSD. This allows running DeepSeek-V4-Flash 284B with 5.3GB practical memory usage and 4.8 tok/s on 24GB Mac hardware.
HOW THIS AFFECTS YOU
●
builderYou can now deploy massive MoE models on consumer-grade hardware using SSD-based expert streaming.