Large Language Model (LLM) inference is increasingly served on disaggregated clusters that combine different accelerator types, including NVIDIA GPUs, Huawei Ascend NPUs, and Kunlunxin AI processors. We show that these heterogeneous xPUs exhibit distinct energy-performance behaviors across inference phases and model mo...
Dong Dong, Cheng-Zhang Wu, Zheng Chen et al.· IEEE Transactions on Paralle...· 0 citations
Trace-based performance analysis provides essential insights for understanding and optimizing large-scale parallel applications. However, traces from applications running on tens of thousands of processes can easily exceed terabytes, far surpassing the memory capacity of typical computing nodes. Existing approaches eit...
Yu-Yang Jin, Ji-Dong Zhai· Proceedings of the Internati...· 0 citations
Trace-based performance analysis provides essential insights for understanding and optimizing large-scale parallel applications. However, traces from applications running on tens of thousands of processes can easily exceed terabytes, far surpassing the memory capacity of typical computing nodes. Existing approaches eit...
Yuyang Jin, Ji-Dong Zhai· Proceedings of the Internati...· 0 citations
A serving system that treats the diffusion block as a compilation unit that improves end-to-end execution time by up to 2.7 times over the strongest surviving baseline under the same 8-GPU placement and remains feasible at the largest batch sizes where multiple baselines run out of memory, while preserving task quality...
Jia-Nian Zhu, Hang Wu, Ying-Hui Li et al.· 0 citations
UniEP fuses the MoE communication and computation into MegaKernels, effectively transforming complex architectural tuning into a unified parameter search space for automated adaptability and incorporates a deterministic token ordering mechanism that guarantees numerical consistency with sequential execution, even under...
Size Zheng, Xuegui Zheng, Li-Wen Chang et al.· IEEE International Symposium...· 1 citation
Prism abstracts the highly dynamic diffusion workload into a predictable, static execution flow transparent to the compiler, and achieves this via three techniques: spatial regularization, temporal stabilization, and specialized kernels that selectively bypass padding data.
Jianian Zhu, Hang Wu, Yinghui Li et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.