Skip to content
Open access

Scalable AI Inference Pipelines Across Edge and Cloud Computing Environments

Aug 2026 · International Journal of Computer Science & Information System · Vol 11, pp. 84-92 · 0 citations

TL;DR

A scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources is examined, providing a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.

Abstract

The increasing deployment of artificial intelligence (AI) applications in healthcare, industrial Internet of Things (IIoT), intelligent transportation, and next-generation wireless systems has created a demand for inference architectures that simultaneously provide low latency, scalability, privacy, reliability, and efficient resource utilization. Conventional cloud-centric inference architectures provide substantial computational capacity but can introduce network latency, bandwidth consumption, privacy exposure, and dependence on centralized infrastructure. Edge computing addresses several of these limitations by relocating computation closer to data sources, while cloud environments remain important for computationally intensive and globally coordinated workloads. This research examines a scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources. The methodology synthesizes the provided literature on federated learning, edge resource allocation, dynamic scheduling, privacy preservation, machine learning for 6G, IIoT, and secure healthcare systems. A layered architectural model is developed around workload characterization, adaptive task placement, communication-aware scheduling, privacy protection, and resilient orchestration. The analysis indicates that scalability is not achieved merely by adding computational resources; rather, it depends on coordinated optimization of computation, communication, privacy, and scheduling. The proposed conceptual framework positions edge inference as the first computational layer, cloud inference as an elastic computational layer, and intelligent orchestration as the mechanism connecting the two. The resulting architecture provides a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.

Read PDF

Similar papers

Review Open access Aug 2026

Edge-Cloud AI Computing: A Robust Framework for Real-Time Inference and Decision Automation

Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments.

Doni Setiawan, Ditha Permata · 0 citations
Open access Aug 2026

Resilient Edge-to-Cloud AI Architectures for Distributed Real-Time Decision Making

A research-driven conceptual framework for resilient edge-to-cloud AI architectures supporting distributed real-time decision making and identifies limitations associated with heterogeneous devices, uncertain ground truth, model drift, communication failures, and the absence of uniform evaluation criteria are identified.

Chinedu Eze, Fatima Bello · 0 citations
Conference Open access Jul 2026

Profiling Neural Network Partitioning Strategies for Inference across the Computing Continuum

As deep learning permeates latency-sensitive domains such as autonomous driving and smart surveillance, deploying neural networks (NNs) across the computing continuum (CC), from IoT devices to edge servers and cloud platforms, has become increasingly important. In such heterogeneous IoT-Edge-Cloud environments, distributed inference promises reduced latency, improved privacy, and better resource utilization. Yet, determining how to deploy NNs over heterogeneous IoT-Edge-Cloud nodes remains a difficult and largely manual process. This paper presents a principled and extensible framework for evaluating distributed inference of NNs in heterogeneous CC infrastructures. We introduce a formal model that unifies functional, pipelined, and data-parallel partitioning strategies within a single abstraction over heterogeneous CC topologies, enabling structured cross-strategy comparison. Building on this foundation, we implement a distributed inference orchestrator that supports flexible deployment of partitioned CNNs, and introduce PartiBench, a benchmarking tool that profiles segments and guides their placement. Our evaluation demonstrates how the framework exposes key performance trade-offs, offering actionable insights into latency, memory use, and communication overhead across IoT-Edge-Cloud nodes. These contributions enable empirical, cross-strategy comparison of distributed inference deployments and provide a basis for future automated placement methods in heterogeneous IoT-Edge-Cloud systems.

Nikolaos Papadakis, Alexandros Angourakis, K. Magoutis et al. · 0 citations
Open access 2020

Hybrid Cloud-Edge Infrastructures for Scalable IIoT AI Deployments

This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.

Jennifer Clark · 0 citations
Conference Jul 2026

Novel Hierarchical Edge AI Architecture for Resource-Constrained Embedded Platforms: A Comprehensive Framework for Distributed Intelligence

This paper presents a comprehensive analysis of edge AI architectures targeting embedded platforms and proposes a novel hierarchical design that addresses the critical challenges of computational efficiency, power consumption, and real-time processing in resource-constrained environments. The proposed architecture integrates adaptive quantization, dynamic load balancing, and multi-tier processing to optimize AI inference at the edge while maintaining high accuracy and low latency. Current edge AI implementations, such as ESP32-based systems, demonstrate the feasibility of bringing artificial intelligence to embedded devices, but lack the sophisticated resource management and scalability required for complex AI workloads. Our literature review reveals significant gaps in existing architectures, particularly in handling dynamic workloads and optimizing resource utilization across heterogeneous computing elements. We propose a three-tier hierarchical edge AI framework that couples adaptive mixedprecision quantization with a cross-tier load balancer and monitoring place, allowing the system to dynamically choose both precision and execution tier based on energy, latency, and accuracy constraints

Ravi Suppiah, Ravichandran Danthakanni, M. Nair et al. · 0 citations
Review Aug 2026

A Intelligent Edge-Cloud Integration for Resilient and Real-Time AI Decision Systems

The analysis indicates that effective edge-cloud AI systems require adaptive workload placement, privacy-preserving distributed learning, security-aware inference, explainability, fault tolerance, and continuous resource optimization rather than simple physical distribution of computation.

Dr. Amir Hosseini, dr.nematollah karimi · 0 citations