Converting dense models to sparse MoE without retraining
October 11, 2026
A method for upcycling dense models like Qwen2.5-0.5B and SmolLM2-360M into sparse Mixture-of-Experts architectures without pretraining from scratch. Using 32 experts with top-8 routing, these conversions aim to reduce per-token inference costs while maintaining existing knowledge density.
HOW THIS AFFECTS YOU
●
builderYou can potentially lower inference latency and costs by converting small dense models into MoE versions.
●
researcherThis explores the limits of sparse upcycling on consumer-grade hardware like an RTX 4060.