RAZOR: Training-Free Expert Pruning via Consensus Residuals
September 28, 2026
RAZOR enables training-free Mixture-of-Experts (MoE) pruning by scoring expert replaceability using consensus residuals. By measuring deviations of expert outputs from the original weighted mixture, it reduces storage on models like Qwen3.6-35B-A3B and DeepSeek-V4-Flash without requiring gradients or recovery training.
HOW THIS AFFECTS YOU
●
builderYou can reduce MoE model storage requirements without the high cost of retraining.
●
researcherThe method provides a new way to evaluate expert contribution through functional replaceability.