Sep 2026· Proceedings of the International Conference on Parallel Processing· pp. 715-725· 1 citation· 14 references
Abstract
In cloud data centers, the colocation of multiple services and applications on the same physical server is crucial for maximizing resource utilization and reducing utility costs. Unfortunately, contention for shared resources across cores, such as the Last-Level Cache (LLC), may lead to severe interference among applications. This complicates quality-of-service (QoS) enforcement and causes performance disparities among applications, ultimately degrading user experience and system fairness. To address these issues, this paper proposes RDPart, an OS-level Reuse-Driven LLC Partitioning policy designed to improve fairness while preserving the QoS of cloud workloads. In contrast to existing fair LLC-partitioning techniques, RDPart eliminates the need for online performance profiling under varying LLC allocations, thereby avoiding undesired QoS violations typically stemming from such invasive exploratory approaches. Furthermore, RDPart adopts a black-box design, making it well-suited for public cloud environments where real-time QoS feedback from applications is unavailable. We implement RDPart in the Linux kernel and evaluate its effectiveness across a diverse set of workloads, including cloud services and HPC/scientific applications. The results reveal that RDPart achieves a 32.1% average fairness improvement with respect to a state-of-the-art black-box cache-partitioning proposal, while keeping QoS violations consistently low across workloads.
The results show that soft SLO limits reduce corrective rescheduling actions by 49% compared to hard-limit approaches while maintaining acceptable performance guarantees, and resource-aware scheduling decreases node-level congestion and further mitigates SLO violations, demonstrating the effectiveness of incorporating...
Oliver Larsson, Thijs Metsch, Cristian Klein et al.· 0 citations
With the widespread adoption of cloud computing and virtualization, multi-tenant architecture has become the mainstream deployment model for data center storage. Solid-state drives (SSDs), leveraging high IOPS, low latency, and high parallelism, serve as the core storage medium. Nevertheless, internal resource contenti...
Dan-Dan Su, Qi-Hao Liu, Xin-Ming Li et al.· International Conference on...· 0 citations
Large language models (LLMs) increasingly rely on context caching to enhance serving efficiency. However, this optimization inadvertently compromises fairness in multi-tenant LLM serving systems. Existing fair schedulers, which account only for compute resources, are unable to handle the multi-dimensional resource dema...
Zhuo-Yan Bai, Bin Gao, Fei Xu et al.· IEEE Transactions on Paralle...· 0 citations
Modern cloud platforms provide diverse compute services, including virtual machines, Function-as-a-Service, and Query-as-a-Service, each offering unique trade-offs in performance, elasticity, and cost. While these services collectively cover the diverse needs of OLAP workloads, existing systems typically rely on a sing...
Wen-Bo Li, Hao-Qiong Bian, Chao Zhang et al.· Proceedings of the ACM on Ma...· 0 citations
In cloud data centers, resource overcommitment is a key strategy for improving energy efficiency, yet managing it to prevent Service Level Agreement (SLA) violations remains a persistent challenge. Virtual machine replacement (VMrP)—the periodic redistribution of running VMs across servers—is central to this ongoing ma...
Hyeongbin Kang, Hyeon-Jin Yu, Heeju Kang et al.· Journal of Cloud Computing· 0 citations
FairInference provides the novelelta-token fairness guarantee: for a well-behaved client, if a token is generated in d time units in isolation, it will be generated within d + {\delta} time units in multi-tenant execution, providing strong latency isolation guarantees for LLM serving.
Dev Bali, Soujanya Ponnapalli, Yi-Chu Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.