Aug 2026· International Journal of Computer Science & Information System· 0 citations
TL;DR
Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments.
Abstract
The rapid deployment of artificial intelligence (AI) across distributed environments has created a need for computing architectures capable of simultaneously supporting low-latency inference, scalable model execution, communication efficiency, privacy, and reliable decision automation. Conventional cloud-centric AI architectures provide substantial computational capacity but can introduce communication delays, bandwidth dependency, privacy exposure, and service disruption risks when real-time decisions must be generated close to data sources. Edge-cloud AI computing addresses these limitations by distributing inference, coordination, and computational workloads across edge devices and cloud infrastructure. This research and review article develops a robust conceptual framework for real-time edge-cloud AI inference and decision automation by synthesizing research on federated learning, communication compression, heterogeneous model aggregation, blockchain-enabled healthcare systems, reinforcement learning, and distributed AI pipelines. The analysis identifies communication efficiency, heterogeneous computational capabilities, privacy preservation, adaptive resource allocation, and resilient inference coordination as the principal architectural requirements. The proposed framework integrates edge-level preprocessing and inference, adaptive communication, cloud-level model coordination, and automated decision orchestration. Findings indicate that compression and knowledge distillation can reduce communication burdens, while heterogeneous aggregation and adaptive learning mechanisms improve the practicality of distributed AI environments. Reinforcement learning further provides a mechanism for dynamic resource and decision optimization. However, the framework remains constrained by device heterogeneity, synchronization overhead, model inconsistency, security requirements, and the trade-off between inference accuracy and latency. The study establishes an integrated theoretical foundation for designing scalable and resilient edge-cloud AI systems capable of supporting real-time intelligent decision processes.
A scalable edge-to-cloud AI inference pipeline in which inference tasks are dynamically distributed across heterogeneous edge and cloud resources is examined, providing a basis for resilient real-time AI systems while highlighting unresolved challenges involving heterogeneous hardware, dynamic workloads, privacy-utility trade-offs, and cross-layer optimization.
Dr. Khalid Al- Mansour· International Journal of Com...· 0 citations
A research-driven conceptual framework for resilient edge-to-cloud AI architectures supporting distributed real-time decision making and identifies limitations associated with heterogeneous devices, uncertain ground truth, model drift, communication failures, and the absence of uniform evaluation criteria are identified.
Chinedu Eze, Fatima Bello· International Journal of Adv...· 0 citations
The analysis indicates that effective edge-cloud AI systems require adaptive workload placement, privacy-preserving distributed learning, security-aware inference, explainability, fault tolerance, and continuous resource optimization rather than simple physical distribution of computation.
Dr. Amir Hosseini, dr.nematollah karimi· International Journal of Adv...· 0 citations
Real-time monitoring and smart decision-making are required in cyber-physical infrastructures such as smart grids, transportation systems, and industrial automation systems to ensure process efficiency and system resilience. However, higher latency, bandwidth constraints, and a lack of responsiveness to time-sensitive data streams haunt traditional cloud-based architectures. To address these challenges, a coherent framework for integrating AI-based cloud analytics with edge intelligent hardware for managing cyber-physical infrastructure is proposed in the following paper. The architecture is based on distributed edge nodes, with hardware accelerators for very low-latency inference, and cloud layers that perform large-scale analytics and optimisation of global models. A hybrid resource allocation strategy is a self-regulating plan for the allocation of computing workloads across cloud and edge environments. Also, an adaptive learning mechanism improves prediction accuracy in changing operational environments. The framework is mathematically designed to achieve optimal latency, throughput, and computational efficiency. The proposed system reduces latency by 145ms to 92 (≈36.5) units and increases prediction accuracy from 81.2% to 96.4% when applied to a dynamic workload. The convergence analysis shows that standardized quicker convergence occurs after 35 iterations as opposed to 60 iterations in the models at the baseline. Additionally, the throughput is proceeding at 520 requests to 780 requests and error rates are falling by a factor of around 41, ensuring better reliability. Relative performance analysis across various scenarios indicates consistent improvement in both edge-dominant and cloud-dominant setups. The results of this study demonstrate that, when combined with hardware-based edge intelligence, AI-powered cloud analytics can significantly enhance the responsiveness, scalability, and decision accuracy in cyber-physical infrastructure systems.
Naveen, Satyam Kumar Sainy· International Journal on Eng...· 0 citations
This paper explores hybrid cloud-edge infrastructures as a scalable solution for deploying AI in IIoT environments and presents an architectural framework that balances compute-intensive model training in the cloud with low-latency inference at the edge.
Jennifer Clark· International Journal of Mac...· 0 citations
This paper presents a comprehensive analysis of edge AI architectures targeting embedded platforms and proposes a novel hierarchical design that addresses the critical challenges of computational efficiency, power consumption, and real-time processing in resource-constrained environments. The proposed architecture integrates adaptive quantization, dynamic load balancing, and multi-tier processing to optimize AI inference at the edge while maintaining high accuracy and low latency. Current edge AI implementations, such as ESP32-based systems, demonstrate the feasibility of bringing artificial intelligence to embedded devices, but lack the sophisticated resource management and scalability required for complex AI workloads. Our literature review reveals significant gaps in existing architectures, particularly in handling dynamic workloads and optimizing resource utilization across heterogeneous computing elements. We propose a three-tier hierarchical edge AI framework that couples adaptive mixedprecision quantization with a cross-tier load balancer and monitoring place, allowing the system to dynamically choose both precision and execution tier based on energy, latency, and accuracy constraints
Ravi Suppiah, Ravichandran Danthakanni, M. Nair et al.· 2026 6th International Confe...· 0 citations