This work introduces MARS (Monte Carlo Tree Search-based Adaptive and Responsive Scheduler), a training-free HPC scheduler whose optimization goal is configurable through a reward function rather than baked into a learned model.
Abstract
Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management. Deep Reinforcement Learning (DRL) has shown promising scheduling performance but requires historical training data and fixes the optimization goal at training time, forcing operators to retrain whenever priorities shift. We introduce MARS (Monte Carlo Tree Search-based Adaptive and Responsive Scheduler), a training-free HPC scheduler whose optimization goal is configurable through a reward function rather than baked into a learned model. MARS uses a lightweight discrete-event simulator to explore the future consequences of scheduling decisions within a strict time budget, adapting to the configured reward at each scheduling cycle. We evaluate MARS on year-long production workloads from two systems at Argonne Leadership Computing Facility -- 4,360-node Theta and 560-node Polaris---under two reward functions: wait-time minimization (MARS-CW) and utilization maximization (MARS-CU). Unlike DRL and heuristics, which only react to the current queue or wait for backfill to find holes, MARS exploits look-ahead to proactively drain the system and plan around future reservations, packing the system to avoid the fragmentation and utilization drop that typically precede reservation windows. MARS-CW reduces tail wait time by 64% on Theta and 43% on Polaris over the production WFP heuristic, while MARS-CU recovers utilization in the 48 hours leading into maintenance, demonstrating that MARS can target either objective via reward reconfiguration.
INTRODUCTION: Edge-cloud schedulers must coordinate latency, energy, load balance, and deadline compliance under changing demand while keeping task-arrival and throughput units physically consistent.OBJECTIVE: This study evaluates MORL-ECSO under an auditable, paired-seed simulation protocol and compares it with tuned...
Li-Na Guo, Cheng-Yu Sun· ICST Transactions on Scalabl...· 0 citations
Job scheduling in Electronic Design Automation (EDA) environments presents unique challenges due to high-frequency job submissions, short job durations, and strict latency requirements. Production schedulers such as IBM Spectrum LSF employ robust heuristics like First-Come-First-Served (FCFS) that provide predictable b...
Yiming Shao, Aijun An, Michael Spriggs et al.· Proceedings of the 32nd ACM...· 0 citations
Cloud computing resource allocation remains a critical challenge, with organizations wasting an estimated $109 billion annually on idle or over-provisioned resources. Traditional allocation strategies—static provisioning, threshold-based autoscaling, and time-series forecasting—fail to capture the complex, non-stationa...
Msr Prasad· International Journal of Tec...· 0 citations
Power-state management in high-performance computing (HPC) clusters must reduce idle energy without excessive wake-up delays for rigid parallel jobs. This paper presents SNF-ICON, an event-driven controller combining smallest-need-first (SNF) gang scheduling, predictive wake timing, and adaptive warm-spare control. At...
BOOSTEDSOSA is introduced, a dual-FPGA ML-assisted Scheduling architecture that integrates a Machine Learning predictor for expected processing times, with a novel temporal-aware training policy, enabling its use in existing HPC systems.
Adam H. Ross, Riccardo Revalor, Aryan Singh et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.