As applications demand increasing memory capacity in clouds, memory pooling provides a cost-effective way to improve utilization and expand capacity. Compute Express Link (CXL), which enables high-performance direct access to remote memory, makes this approach increasingly practical. However, existing studies of memory...
Guang-Qiang Luan, Pu Pang, Quan Chen et al.· Proceedings of the Internati...· 0 citations
MEDO leverages a novel multi-stream data offloading architecture, featuring parallel data streams and approximate LRU queues, to maximize throughput and efficiently handle diverse workloads and incorporates a lightweight, adaptive offloading agent that dynamically optimizes data placement decisions and fine-grained sys...
Jing Wang, Han-Zhang Yang, Chao Li et al.· Proceedings of the Internati...· 0 citations
The scale of data-intensive workloads in intelligent datacenters has grown rapidly in recent years, intensifying the need for efficient data movement within disaggregated memory pools across memory and storage subsystems. However, most existing solutions focus primarily on optimizing communication between compute nodes...
Jing Wang, Han-Zhang Yang, Chao Li et al.· Proceedings of the Internati...· 0 citations
As applications demand increasing memory capacity in clouds, memory pooling provides a cost-effective way to improve utilization and expand capacity. Compute Express Link (CXL), which enables high-performance direct access to remote memory, makes this approach increasingly practical. However, existing studies of memory...
Guangqiang Luan, Pu Pang, Quan Chen et al.· Proceedings of the Internati...· 0 citations
This work proposes a series of optimizations for these two kernels, including computation-transfer pipelining, load balancing, and memory access fusion, achieving 1.97 × to 2.16 × proof generation speedup over a state-of-the-art open source GPU acceleration library.
Xinwei Qiang, Liukun Yu, Xiyu Wang et al.· IEEE International Symposium...· 0 citations
The increasing use of renewable energy in data centers creates an opportunity to reduce the carbon footprint of energy-intensive LLM inference workloads. Unlike traditional stable power supply, renewable generation fluctuates over time, making it difficult to match computation demand with available energy. However, exi...
Chang Liu, Jiacheng Liu, Xiaofeng Hou et al.· Fall Joint Computer Conferen...· 0 citations
Prefix caching has become a key technique for LLM serving, and nowadays the reusable KVCache contents are often hosted on distributed servers. For long-context LLM inferences with high cache hit ratio, cross-server KVCache transmission has become an emerging performance bottleneck; such network-intensive LLM inferences...
Weiye Wang, Chen Chen, Junxue Zhang et al.· Asia-Pacific Workshop on Net...· 0 citations
Federated Learning (FL) allows edge clients to collaborate in model training with data privacy preserved, yet it is known to suffer low training efficiency and model accuracy. Given that efficiency and accuracy are usually conflicting objectives, existing practices increasingly employ an adaptive scheme that changes th...
Jiayi Zhang, Zuo Gan, Chen Chen et al.· IEEE Transactions on Mobile...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.