A Formally Verified AI Framework for Scheduling OptimiSation in Slurm-Based Distributed HPC Systems
Abstract
This paper presents an artificial intelligence (AI)-enabled scheduling framework for Slurm that integrates formal verification techniques with data-driven optimization methods. The proposed architecture combines concepts from queueing theory, Markov chain modeling, graph theory, mathematical optimization, reinforcement learning, and swarm intelligence to support adaptive scheduling decisions while preserving system correctness and reliability. By unifying these complementary approaches, the framework enhances scheduling efficiency, improves resource utilization, and increases operational robustness in dynamic HPC environments. Experimental evaluation demonstrates that the proposed framework consistently outperforms conventional scheduling ap-proaches across multiple performance metrics. Specifically, it achieves higher job throughput, shorter queue waiting times, improved resource utilization, and stronger fault tolerance, highlighting its effectiveness for next-generation HPC workload management.