Aug 2026· ACM Transactions on Architecture and Code Optimization (TACO)· 0 citations· 22 references
TL;DR
Experimental results across diverse HPC and AI workloads demonstrate that HARMONI significantly reduces average makespan by up to 90% compared to state-of-the-art schedulers, effectively bridging the gap between evolving storage hardware and the dynamic I/O demands of modern HPC systems.
Abstract
Modern HPC systems increasingly rely on tiered storage architectures with SSDs serving as a critical performance tier. However, the inherent asynchronous I/O characteristics of SSDs, including read/write bandwidth asymmetry and interference, pose significant challenges for traditional I/O schedulers. These challenges are exacerbated by the convergence of bursty HPC write workloads (e.g., checkpointing) and sustained AI read workloads (e.g., data streaming) on shared SSD infrastructure. Existing schedulers fail to adequately address these combined workloads, leading to suboptimal resource utilization. This paper introduces HARMONI, a heterogeneity-aware reinforcement learning scheduler for mixed I/O in HPC storage systems. HARMONI leverages a graph neural network (GNN) to encode task-SSD dependencies and a hybrid interference predictor to adapt to hardware and I/O variations. Experimental results across diverse HPC and AI workloads demonstrate that HARMONI significantly reduces average makespan by up to 90% compared to state-of-the-art schedulers, effectively bridging the gap between evolving storage hardware and the dynamic I/O demands of modern HPC systems.
I/O schedulers are the key component to maintain high quality of services under homogeneous hybrid storage workloads, e.g., multiple concurrently running sustained workloads. However, the long hardware queue for I/O requests under I/O schedulers, which is utilized to exploit the substantial parallelism of underlying st...
Yong-Jun Pan, Le Yu, Jia-Lin Liu et al.· ACM Transactions on Architec...· 0 citations
With the widespread adoption of cloud computing and virtualization, multi-tenant architecture has become the mainstream deployment model for data center storage. Solid-state drives (SSDs), leveraging high IOPS, low latency, and high parallelism, serve as the core storage medium. Nevertheless, internal resource contenti...
Dan-Dan Su, Qi-Hao Liu, Xin-Ming Li et al.· International Conference on...· 0 citations
This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...
Hanqing Li, Tie-Jun Li, Sheng Ma et al.· ACM Transactions on Design A...· 0 citations
Ultra-low-latency (ULL) SSDs enable cloud service providers to co-locate latency-sensitive services and throughputoriented background jobs on the same machines. However, their microsecond-scale latency creates a scheduling dilemma: conventional I/O schedulers introduce visible overhead, while disabling scheduling remov...
Ming Wei, Tzu-Chieh Huang, Chieh-Lin Tsai et al.· IEEE International Conferenc...· 0 citations
Concord, a novel GPU sharing-enabled workload scheduler that outperforms state-of-the-art schedulers, achieves a 1.68 × reduction in JCT and a 29% improvement in GPU utilization in high-load scenarios.
Xin-Hua Wang, Wei-Wei Lin, Hai-Jie Wu et al.· 0 citations
MEDO leverages a novel multi-stream data offloading architecture, featuring parallel data streams and approximate LRU queues, to maximize throughput and efficiently handle diverse workloads and incorporates a lightweight, adaptive offloading agent that dynamically optimizes data placement decisions and fine-grained sys...
Jing Wang, Han-Zhang Yang, Chao Li et al.· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.