Skip to content
Open access

HARMONI: Heterogeneity-Aware I/O Scheduling for Mixed Workloads in SSD-Based HPC Systems

Aug 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · 0 citations · 22 references

TL;DR

Experimental results across diverse HPC and AI workloads demonstrate that HARMONI significantly reduces average makespan by up to 90% compared to state-of-the-art schedulers, effectively bridging the gap between evolving storage hardware and the dynamic I/O demands of modern HPC systems.

Abstract

Modern HPC systems increasingly rely on tiered storage architectures with SSDs serving as a critical performance tier. However, the inherent asynchronous I/O characteristics of SSDs, including read/write bandwidth asymmetry and interference, pose significant challenges for traditional I/O schedulers. These challenges are exacerbated by the convergence of bursty HPC write workloads (e.g., checkpointing) and sustained AI read workloads (e.g., data streaming) on shared SSD infrastructure. Existing schedulers fail to adequately address these combined workloads, leading to suboptimal resource utilization. This paper introduces HARMONI, a heterogeneity-aware reinforcement learning scheduler for mixed I/O in HPC storage systems. HARMONI leverages a graph neural network (GNN) to encode task-SSD dependencies and a hybrid interference predictor to adapt to hardware and I/O variations. Experimental results across diverse HPC and AI workloads demonstrate that HARMONI significantly reduces average makespan by up to 90% compared to state-of-the-art schedulers, effectively bridging the gap between evolving storage hardware and the dynamic I/O demands of modern HPC systems.

Read PDF

Similar papers

Open access Sep 2026

Boosting Burst I/O Performance under Heterogeneous Hybrid Storage Workloads via DUal I/O IssuE Thresholds

I/O schedulers are the key component to maintain high quality of services under homogeneous hybrid storage workloads, e.g., multiple concurrently running sustained workloads. However, the long hardware queue for I/O requests under I/O schedulers, which is utilized to exploit the substantial parallelism of underlying st...

Yong-Jun Pan, Le Yu, Jia-Lin Liu et al. · 0 citations
Conference Aug 2026

QoS Guarantees and Performance Isolation for Multi-tenant SSDs in Cloud Environments

With the widespread adoption of cloud computing and virtualization, multi-tenant architecture has become the mainstream deployment model for data center storage. Solid-state drives (SSDs), leveraging high IOPS, low latency, and high parallelism, serve as the core storage medium. Nevertheless, internal resource contenti...

Dan-Dan Su, Qi-Hao Liu, Xin-Ming Li et al. · 0 citations
Open access Aug 2026

SAI: Virtualizing Shared Memory of GPU for AI workload acceleration

This work proposes SAI, a mechanism that virtualizes shared memory into the L2 cache to improve GPU performance for AI applications and introduces an L2 cache management strategy that integrates associativity-based virtual page allocation and a replacement information table, reducing page-swapping overhead while preser...

Hanqing Li, Tie-Jun Li, Sheng Ma et al. · 0 citations
Conference Aug 2026

ARDA: I/O Scheduler for Heterogeneous Workloads Co-located on Ultra-low-latency SSDs

Ultra-low-latency (ULL) SSDs enable cloud service providers to co-locate latency-sensitive services and throughputoriented background jobs on the same machines. However, their microsecond-scale latency creates a scheduling dilemma: conventional I/O schedulers introduce visible overhead, while disabling scheduling remov...

Ming Wei, Tzu-Chieh Huang, Chieh-Lin Tsai et al. · 0 citations

Future Generation Computer Systems

Concord, a novel GPU sharing-enabled workload scheduler that outperforms state-of-the-art schedulers, achieves a 1.68 × reduction in JCT and a 29% improvement in GPU utilization in high-load scenarios.

Xin-Hua Wang, Wei-Wei Lin, Hai-Jie Wu et al. · 0 citations
Book Open access Sep 2026

MEDO: Adaptive Multi-Stream Data Offloading for Efficient Management of Disaggregated Memory Pool

MEDO leverages a novel multi-stream data offloading architecture, featuring parallel data streams and approximate LRU queues, to maximize throughput and efficiently handle diverse workloads and incorporates a lightweight, adaptive offloading agent that dynamically optimizes data placement decisions and fine-grained sys...

Jing Wang, Han-Zhang Yang, Chao Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.