Optimizing Cloud Resource Management: AI Approaches with Comparative Analysis
Abstract
The fast evolution of cloud environments has created more dynamic, diverse, and latency-sensitive workloads that require greater sophistication in resource management. Traditional, reactive heuristic scheduling methods tend to overlook shifting workloads, resulting in poor resource management, higher energy consumption, and repeated SLA breaches. The aim of this survey paper is to facilitate the use of Artificial Intelligence (AI) in the area of intelligent, proactive, adaptive, and cost-effective scheduling in cloud and edge-cloud infrastructures. This paper performs an extensive analysis of the literature on the use of hybrid and hierarchical deep learning architectures, reinforcement learning techniques (RL), swarm-based methods, hybrid metaheuristic optimization and costsensitive scheduling methods. Each of these techniques is analysed in the light of key evaluation criteria: prediction quality, computational complexity, and real-time scalability, as well as energy consumption. Beyond a literature review, this research presents a benchmark case study on a selected set of models including Linear Regression, Random Forest, and Gradient Boosting, with a focus on model prediction error when using benchmark workload trace datasets. Study results show that of the models studied, the Random Forest model is by far the best, reaching an R2 value of 0.999 in the Alibaba dataset, with the least amount of error.