Toward an Integrated Theory of Adaptive Scheduling in High Performance Computing: A Queuing-Theoretic and Computational Learning Perspective.
Abstract
High Performance Computing (HPC) systems increasingly operate under heterogeneous workloads, dynamic resource availability, and stringent performance and energy constraints. Traditional batch scheduling policies such as First-Come First-Served (FCFS), backfilling, and priority-based heuristics rely on static assumptions about job behavior and system state, often leading to suboptimal utilization and long waiting times in highly variable environments. This paper proposes an integrated theoretical framework for adaptive HPC scheduling that unifies queuing-theoretic models with computational learning techniques. By interpreting job arrivals and service processes through stochastic queues while enabling scheduling decisions to evolve via data-driven learning, we establish a principled basis for adaptive schedulers that can respond to workload uncertainty. We outline the mathematical foundations of this approach, discuss learning-augmented scheduling policies, and present illustrative scenarios demonstrating how adaptive strategies can outperform static heuristics in terms of mean response time, fairness, and system utilization. This work aims to bridge the gap between analytical scheduling theory and practical intelligent resource management in HPC systems.