SlimWise: Decoupling Expert Pruning for Efficient MoE Serving
September 27, 2026
SlimWise is a serving framework that optimizes Mixture-of-Experts (MoE) inference by using the full model during prefill and a pruned model during decoding. This training-free approach allows for a seamless KV cache handoff, minimizing the throughput bottleneck of expert-weight traffic without sacrificing significant accuracy.
HOW THIS AFFECTS YOU
●
builderYou can improve MoE inference throughput by tailoring expert activation to the specific requirements of prefill and decode phases.