TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

arXiv CS.AI
AI Hardware

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. A max-affine profile $t=\max(a+bG,\,c+\beta N)$ captures both regimes.

Related Articles