Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency...
Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al.· Proceedings of the Internati...· 0 citations
Training large-scale graphs with GNNs on multi-GPU platforms faces substantial feature loading overhead, leading to low resource utilization and inefficient training. Overcoming such communication bottlenecks is crucial, and current solutions achieve this by overlapping computation and communication through pipelining....
Jia-Qi Si, De-Zun Dong· ACM Transactions on Architec...· 0 citations
Driven by cost and privacy constraints, many organizations deploy large language model (LLM) inference on local CPU-GPU servers. For long-context requests, the key-value (KV) cache grows linearly with sequence length and can rapidly exhaust GPU memory, reducing both the maximum supported context and request concurrency...
Jun-Wen Zhang, Wei-Ling Yang, Jian-Bin Fang et al.· Proceedings of the Internati...· 0 citations
DPIO is presented, a unified I/O processing stack designed to harmonize the collaboration between CPU and DPU in NVMeoF environments, achieving near-optimal system performance across diverse workloads.
Wenhao Gu, Xuchao Xie, Yujuan Tan et al.· IEEE International Symposium...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.