Innovative Node Autoscaling Mechanism in the Kubernetes Ecosystem
Abstract
In this paper, we examine the Cloud-HPC continuum from the perspective of resource allocation and the execution of indistinguishable HPC and Cloud workloads. After briefly analyzing Kubernetes, an open-source platform for orchestrating containerized applications, we identified inefficiencies in resource usage resulting from the continuous activation of all compute nodes. Next, we review existing reactive autoscaling approaches, which can adjust the number of Kubernetes nodes in real time but cannot anticipate future resource needs. To address this limitation, the paper introduces a combination of proactive and reactive AI-driven autoscaling methodologies that predict the number of compute nodes required by analyzing node behavior during workflow execution. This approach optimizes resource utilization, reduces operational costs, and maintains application performance in cloud environments. It integrates seamlessly into the Kubernetes ecosystem with minimal modifications. Furthermore, this study challenges the traditional notion of the HPC–cloud continuum by exploring an extreme hypothesis: replacing the HPC scheduler with a cloud orchestrator. The proposed autoscaling module can serve as a new computing node scheduler, integrated with the default Kubernetes scheduler and potentially with HPC schedulers. Experimental evaluations across diverse benchmarks confirm their effectiveness in improving efficiency, reducing costs, and sustaining performance, benefiting Kubernetes cluster administrators and end users.