FaultSense: Fault Localization in Large-Scale Mixture-of-Experts Model Serving Infrastructure
While Mixture-of-Experts (MoE) models scale LLM serving across hundreds of GPUs, their reliance on all-to-all communication makes them susceptible to gray failures (e.g., GPU stragglers) which inflate serving latency without explicit errors, complicating fault localization. We present FaultSense, an application-layer a...