MigMoE: Task-Aware Expert Migration for Faster and More Balanced Expert-Parallel MoE Inference
Sparse Mixture-of-Experts (MoE) architectures scale LLM capacity, but serving them with expert parallelism often suffers from the straggler effect caused by skewed token routing and uneven placement of hot experts across GPUs. Existing methods mitigate stragglers by adjusting token-to-expert distributions or by replica...