Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems com...
Yao Fei, Gong-Ming Zhao, Hong-Li Xu et al.· 0 citations
AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e....
Yao Fei, Jin Fang, Si-Ze Zheng et al.· 0 citations
Regarding the computational intensity and stringent latency requirements of modern applications comprised of many dependent subtasks, collaborative edge computing (CEC) emerges as a solution to guarantee the service-level objectives (SLOs) by offloading subtasks to multiple distributed edge servers. However, existing o...
Zi-Xuan He, Yan-Jing Sun, Zhen-Guo Ma et al.· IEEE Transactions on Mobile...· 0 citations
Hestia is proposed, a framework that achieves long-term stable oversubscription through workload aggregation through a smoothing-based method to classify workloads suitable for aggregation according to their periodicity, and an aggregation algorithm to minimize the overall MCV.
Baoqing Wang, Gongming Zhao, Hongli Xu et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.