Skip to content
Review

AI-Enabled Cloud Operations for Predictive Monitoring, Automation, and Service Reliability

Jul 2026 · Global academic journal of economics and business · 0 citations

TL;DR

The study proposes a layered architecture linking logs, metrics, traces and service dependency graphs with predictive analytics, automation policy and reliability dashboards, and gives a maturity roadmap for organisations that want to transition from reactive monitoring to proactive, self-improving operations.

Abstract

Cloud operations have become a strategic reliability function as enterprises move customer journeys, public services, industrial platforms and data products to hybrid and multi cloud environments. Traditional monitoring practices are no longer sufficient as cloud systems generate high-velocity telemetry from containers, serverless functions, networks, databases, identity services and user-facing applications. This review looks at how artificial intelligence for IT operations (AIOps) can improve predictive monitoring, controlled automation and service reliability. The paper uses a structured narrative review approach to synthesise literature, standards and implementation guidance published from 2020 to 2025. The review identifies five operational domains, namely observability engineering, incident prediction, root cause analysis, automated remediation, and governance for trustworthy operations. It argues that it is not possible to improve service reliability just through machine learning models, but also needs good telemetry, topology awareness, safe runbooks, cyber resilience, human validation and continuous feedback from incidents. Saudi Arabia is particularly relevant for AI-enabled cloud operations as Vision 2030 digital transformation, cloud-first adoption, smart city initiatives, financial technology, healthcare platforms and critical digital infrastructure require highly available and resilient services. The study proposes a layered architecture linking logs, metrics, traces and service dependency graphs with predictive analytics, automation policy and reliability dashboards. It also gives a maturity roadmap for organisations that want to transition from reactive monitoring to proactive, self-improving operations. The review concludes that AIOps can reduce mean time to detect and recover, improve capacity planning and support business continuity, provided that organisations treat automation as a governed socio-technical capability rather than a purely technical tool.

View source

Similar papers

Review Open access Aug 2026

From Detection to Verified Action: Operational Readiness for AI-Enabled Cloud Failure Management

Artificial intelligence for information technology operations has progressed from alert correlation and anomaly detection to root-cause analysis, mitigation recommendation, and agents that can invoke infrastructure tools. This progression creates an assurance problem: analytical performance does not establish authority to change a live cloud system. This critical narrative review examines the evidence required before AI output may influence or execute failure-management action in mission-critical cloud infrastructure. Searches through 24 July 2026 covered scholarly databases, major systems venues, standards sources, and authoritative production reports. Forty-four sources were coded by operational task, evidence setting, authority, safety control, rollback, and outcome verification. Production evidence is substantial for detection, triage, diagnosis, and several narrowly bounded mitigation systems, but remains weak for general-purpose autonomous action. Recent agentic AIOps surveys emphasize contracts, bounded tools, canary deployment, and rollback; the unresolved need is an assessment method that separates what a system can infer, what it may do, and what evidence shows that the action remained safe. The proposed Operational Decision-Readiness and Verification framework represents a deployment claim through analytical capability, operational authority, and assurance maturity. Twelve domains use explicit 0-3 evidence anchors, while endpoint validity, evidence integrity, intervention risk, authorization, reversibility, and recovery verification operate as non-compensatory gates. Worked assessments of six systems illustrate distinctions among analytical support, testbed execution, and narrow production autonomy. The framework is conceptual and requires prospective and inter-rater validation; it structures an actionspecific readiness case rather than certifying a model or product

Adepegba Akindayomi Akintade · 0 citations
Review Open access 2024

The Future of Site Reliability Engineering: AI-Driven Observability and Autonomous Operations in Multi-Cloud Environments

SRE has come a long way since its start, from monitoring infrastructure health to reactive events. Modern SRE is an AI-driven observability platform providing real-time visibility into complex distributed systems. As multi-cloud use grows, operational complexity increases and it becomes increasingly complicated to provide dependability, performance and security across a broad range of cloud platforms. Artificial Intelligence (AI), Machine Learning (ML) and AIOps technologies are changing Service Availability and Incident Response (SRE) with intelligent anomaly detection, predictive analytics, automated root-cause investigation and autonomous remediation in response to such. The effort aims at investigating the future of service-oriented architecture (SRE) in the age of AI-based observability and autonomous operations in multi-cloud environments. Through a review of current technologies, industry practice and upcoming trends in research. The objective of this research is to investigate the viability of the application of AI-based solutions to enhance system dependability, minimize operational overhead and optimize incident response efficiency. The study will also address issues of scalability, interoperability and governance. The results indicate that AI and ML in observability systems can deliver automated operational workflows that help enable proactive reliability management, minimize downtime and accelerate decision making. This makes AIOps powered autonomous operations a rapidly growing important enabler for cloud native infrastructures with self repairing systems. This paper describes the convergence of AI with software defined networking (SRE) and strategic implications for enterprises seeking strong, scalable and efficient cloud operations. The results show intelligent automation is becoming more important in shaping the future of cloud-native reliability management and operational excellence.

Srichandra Boosa · 0 citations
Open access Aug 2026

Artificial Intelligence as Effectiveness Enabler of Dynamic Reconfiguration of Systems Architecture in Industry 5.0

Industrial operations increasingly face high-stakes decisions that involve people, data streams, simulations, and control systems. Urgent sessions often require external expertise, retrieval of documents and live telemetry, running what-if simulations, and verifying safety constraints. These scenarios highlight the need for secure interoperability, explainable decision support, and human-in-the-loop control. This paper presents a proposal of a technology-agnostic reference architecture that builds on Industry 4.0 frameworks by incorporating the human-centric, resilient, and sustainable principles of Industry 5.0. Its intelligent layer enables the new approach to human involvement in the process, facilitating meaningful human–machine collaboration. The proposed research provides a practical and conceptual framework for systems engineers, industrial software architects, and operations managers seeking to transition legacy operational plants into human-aligned ecosystems. Its feasibility is evaluated through a simulation-based underground mining testbed, where heterogeneous data sources and communication protocols are integrated into a common operational environment. The proof of concept shows how telemetry, data storage, machine learning models, and operator feedback can be combined to support auditable, explainable, and human-contestable industrial decisions, demonstrating the classification accuracy, remaining useful life forecasting capabilities, and enhanced recommendation precision enabled by iterative operator feedback loops.

Luis Ferreira, Eduardo Gonçalves, G. Putnik et al. · 0 citations
Open access Jul 2026

Artificial Intelligence-Enabled Predictive Decision Support Systems for Smart Enterprise and Industrial Applications

AI-powered predictive systems for decision support are revolutionizing the way that smart enterprises and industrial organizations are analysing data, predicting future conditions, and making operational and strategic decisions. The systems include machine learning, deep learning, predictive analytics, prescriptive analytics, real-time monitoring, and intelligent recommendation systems to enhance decision-making accuracy, efficiency, and responsiveness. They are used in business forecasting, customer and financial analytics, supply-chain and inventory management, predictive maintenance, production optimization, quality control, energy management, workplace safety and asset monitoring. The addition of new technologies like the Internet of Things, Industrial Internet of Things, digital twins, cloud and edge computing, robotics, blockchain and next generation networks further improve system connectivity, scalability and real-time performance. The successful implementation of these steps needs a structured framework for problem identification, data collection, preprocessing, feature engineering, model selection, training, validation, system integration, deployment, and continual monitoring. Despite these progressions, data quality, interoperability, scalability, algorithmic bias, explainability, privacy, cybersecurity, organizational readiness, and regulatory compliance are all important challenges that still need to be addressed. There is still a need for human oversight, especially when dealing with safety-critical and high-impact decisions. It includes the technological foundations, system architecture, implementation processes, enterprise and industrial applications, performance evaluation, governance requirements, and future directions of AI-supported predictive decision support systems. It concludes that the systems that are trustworthy, secure, transparent, sustainable and intelligent are enterprise and industrial operations.

Smitha Rajagopal, R. Karchi, Sanjeevakumar M.Hatture et al. · 0 citations
Review Open access Jul 2026

Risk management frameworks for cloud and AI systems: A comparative review

The rapid growth of cloud computing and artificial intelligence (AI) has transformed enterprise operations, enabling scalability, automation, and advanced predictive capabilities. However, their convergence introduces complex and interdependent risks, including data breaches, service disruptions, regulatory challenges, and AI-specific concerns such as bias, lack of explainability, and ethical implications. This narrative review synthesizes current literature on risk management frameworks for cloud and AI systems to identify key similarities, differences, and integration opportunities. A structured search of IEEE Xplore, Scopus, and Web of Science was conducted to select high-quality peer-reviewed studies based on methodological rigor, framework comprehensiveness, and relevance to cloud-AI environments. Comparative analysis shows that established cloud frameworks such as NIST RMF, ISO/IEC 27001, and the CSA Cloud Controls Matrix effectively address security, operational, and compliance risks. In contrast, AI-focused frameworks, including ISO/IEC TR 24028, NIST AI RMF, and EU AI guidelines, primarily target governance, ethical, and model-specific risks. Despite these advances, significant gaps remain in unified risk management approaches for hybrid cloud–AI systems, particularly in harmonizing security, compliance, and explainability metrics. To address this, this study proposes the Unified Adaptive Cloud Resilience Framework (UACRF), an integrated risk management model that unifies cloud security, operational risk, AI governance, and model lifecycle risks into a single adaptive framework that combines the strengths of both domains, enabling more adaptive, scalable, and ethically aligned risk mitigation. Unlike existing frameworks that treat cloud and AI risks independently, UACRF provides a unified cross-domain architecture for managing hybrid cloud-AI risk environments. This study also offers actionable insights for the development of resilient, trustworthy, and compliant AI-cloud systems in high-stakes environments. Keywords: Cloud Computing, Artificial Intelligence, Risk Management Frameworks, Integrated Risk Mitigation, Hybrid Cloud-AI Systems.

Oluwafemi Oluwagboyega Fabiyi, Solomon Doe Adjaottor · 0 citations
Open access 2025

Explainable AI-Based Predictive Maintenance Framework for Industrial Equipment Reliability

Industry 4.0 has transformed manufacturing through the integration of Industrial IoT (IIoT), cyber-physical systems, cloud computing, and artificial intelligence, making predictive maintenance (PdM) a key strategy for improving equipment reliability. Unlike traditional maintenance, AI-driven PdM analyzes real-time sensor data to predict equipment failures before they occur. However, many AI models operate as black boxes, limiting trust and interpretability. The proposed Explainable AI-Based Predictive Maintenance Framework (XAI-PMF) addresses this challenge by integrating IIoT sensing, intelligent feature engineering, hybrid machine learning (Random Forest, Gradient Boosting, LSTM, and Transformers), and explainability techniques such as SHAP, LIME, and rule extraction. These methods provide transparent fault predictions and maintenance recommendations by highlighting the factors influencing equipment degradation. Continuous learning further enables adaptive model updates as new operational data become available. Overall, the framework improves prediction accuracy, reduces downtime and false alarms, enhances maintenance scheduling, and supports trustworthy, intelligent asset management for next-generation smart factories.

Narendra Karmarkar · 0 citations