[arXiv]score: 0.14
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
August 14, 2026
TEMPO optimizes expert-parallel MoE serving by modeling expert latency as a max-affine function of HBM weight streaming and grouped GEMM compute. This approach addresses non-linear execution times across memory-bound and compute-bound regimes, where traditional token-count balancing fails. Empirical tests show TEMPO's latency modeling outperforms existing proxy dispatchers by up to 1.7x in p95 makespan.
DAILY DIGEST
you don't check 9 sources — we do. one email every morning, read in 2 min. free. unsubscribe anytime. privacy