DwarfStar Inference Engine Enables 100T Parameter MoE on 96GB VRAM
July 26, 2026
DwarfStar uses a predictive expert paging algorithm to run Mixture of Experts models with 100 trillion total parameters and 80 billion active parameters. The engine achieves extreme memory efficiency by streaming experts into 96GB of total VRAM using four consumer-grade 24GB GPUs.
HOW THIS AFFECTS YOU
●
builderYou can run massive MoE models on consumer hardware using predictive paging.
●
researcherThis approach explores new boundaries for memory streaming and expert paging efficiency.