Jul 2026· Annual International Computer Software and Applications Conference· pp. 1677-1680· 0 citations· 20 references
Abstract
CloudSkin aims to create a cognitive cloud-edge continuum platform that optimally exploits heterogeneous resources. Using AI, the platform automatically adapts to system and application behavior, and enables secure and seamless service deployment. Barcelona Supercomputing Center has led the design and development of an AI-based “Learning Plane” to automatically and continuously manage services in the cloudedge continuum to adapt to the dynamic environment. In the project, we enabled basic system models and workload characterization including regressions and time-series models, different levels of smart policies including heuristics and reinforcement learning, and we developed a data-connector agent to leverage those mechanisms towards real-world use cases. In this paper, we report on the technical results and insights we have explored within the project, explain how they are integrated as a learning plane to support CloudSkin use cases, and finally outline new open research lines.
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
Enterprises are increasingly adopting multi-cloud infrastructures to increase service availability, scalability, fault tolerance and vendor independence. Nevertheless, the allocation of resources among heterogeneous cloud providers is a highly critical issue because the workloads are variable, the latency is limited, the SLA compliance issues are present, and the privacy concerns are also central to the centralized scheduling models. Conventional centralized reinforcement schedulers in learning demand complete visibility of workload, which can cause overheads in the communication and expose risks of data disclosure. In this paper, we suggest a Federated Reinforcement Learning Scheduler (FRLS) in privacy-sensitive and adaptive optimization of resources on multi-clouds. The proposed architecture has every cloud node separately train a reinforcement learning-based scheduling agent on local observations of workload. Rather than exchanging raw data, nodes exchange model parameters with a federated aggregation server on an irregular basis and weighted averaging is used to create a global scheduling model. Simulations with synthetic heterogeneous workload traces show that FRLS utilizes its resources 15-18 percent more efficiently, SLA violations 8-10 percent fewer and converge quicker than heuristic and centralized RL schedulers. The framework offers scalable, secure and distributed intelligence on next-generation multi-cloud orchestration systems.
Y. Gaidhani, Vudutha Sravanthi, D. L. Narayana et al.· International Conference Com...· 0 citations
The proposed approach separates the research-infrastructure layer, which exposes and manages distributed resources, from the application layer, where Cyber-Physical workflows are organized according to an Edge-Fog-Cloud pattern in which placement, timing, and data provenance are treated as first-class experimental concerns.
Fabio Orazio Mirto, Giuseppe Tricomi, L. D’Agati et al.· 0 citations
Cloud computing, despite its scalability advantages, may not fully satisfy the low-latency demands of emerging latency-sensitive applications. The cloud–edge continuum addresses this limitation by integrating the responsiveness of edge resources with cloud scalability. Microservice architecture (MSA), characterized by modular, loosely coupled services, aligns effectively with this continuum. However, heterogeneous and dynamic computing resources pose significant challenges to optimal microservice placement. Most existing approaches focus on generating one-time scheduling plans, which are ill-suited to dynamic environments where frequent and lightweight rescheduling actions are required in response to changing system conditions. We propose REACH, a reinforcement learning-based microservice rescheduling framework that enables a sim-to-real deployment pipeline for adapting microservice placement under fluctuating resource availability and performance variations across distributed infrastructures. REACH is integrated with a real Kubernetes-based cloud–edge continuum testbed, with open-source artifacts released for reproducibility.
Xu Bai, Muhammed Tawfiqul Islam, Rajkummar Buyya et al.· IEEE International Conferenc...· 0 citations
The rapid evolution of machine learning (ML) models and the surge in data volumes necessitate scalable and efficient deployment strategies. Cloud-based distributed systems offer on-demand scalability and resource flexibility, making them ideal for real-time ML model deployment and scaling. This paper explores optimization techniques for cloud-based distributed systems to enhance the deployment and scaling of ML models in real-time applications. We examine the integration of distributed systems and ML within cloud environments, focusing on scalable training and inference mechanisms. Key considerations such as task partitioning, communication overhead, fault tolerance, and resource optimization are discussed. Furthermore, we review auto-scaling techniques, highlighting advancements and challenges in dynamically adjusting resources to meet fluctuating demands. The paper also delves into the application of machine learning for cloud resource provisioning, emphasizing dynamic allocation based on real-time usage patterns. By synthesizing current research and practices, this study provides insights into effectively leveraging cloud-based distributed systems for real-time ML model deployment and scaling.
Emma Roberts, William Hughes· International Journal of Art...· 0 citations