While Mixture-of-Experts (MoE) models scale LLM serving across hundreds of GPUs, their reliance on all-to-all communication makes them susceptible to gray failures (e.g., GPU stragglers) which inflate serving latency without explicit errors, complicating fault localization. We present FaultSense, an application-layer a...
Harish S. A., Vignesh S, Ashwin Kurella et al.· Proceedings of the 17th ACM...· 0 citations
Modern AI workloads demand microsecond-scale network reaction times, forcing data centers to offload congestion control algorithms (CCAs) to hardware. Simultaneously, emerging transport standards introduce diverse congestion signals like delay, CSIG, INT, and packet trimming. Understanding hardware-offloaded CCA adapta...
Meet Dadhania, R. K, Saptarshi Samanta et al.· Conference on Applications,...· 0 citations
High-speed programmable data planes provide opportunities to implement data-driven fast reroute systems that quickly adapt to varying network conditions (e.g., congestion, failures) and improve network performance. The core of these systems has packet-processing algorithms running in the data plane that continuously lo...
S. Harish, S. Vignesh, Divya Pathak et al.· IEEE Transactions on Network...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.