In Mixture-of-Experts models with aggressive load-balancing, router probabilities fail as importance signals due to over-dispersed routing. On gpt-oss-20B, low-perplexity pruning configurations yield significantly worse mathematical reasoning performance compared to high-perplexity configurations, unlike standard Mixtral-style routing.
HOW THIS AFFECTS YOU
●
builderBe cautious when using perplexity as a proxy for model quality during MoE compression.
●
researcherYou must account for routing dispersion when evaluating expert importance for pruning.