Aug 2026· International Journal of Innovative Science and Research Technology· 0 citations· 30 references
TL;DR
The findings demonstrate that DRL-driven adaptive orchestration can become a central mechanism for autonomous edge intelligence in next-generation AI-native communication infrastructures.
Abstract
Edge computing has emerged as a foundational paradigm for intelligent digital infrastructure because it reduces
latency, improves bandwidth utilization, and enables real-time analytics close to data sources. Yet modern edge
environments remain highly volatile. Resource availability changes continuously. IoT traffic fluctuates unpredictably.
Mobile users migrate across heterogeneous networks. Conventional heuristic-based schedulers struggle to maintain stable
Quality of Service (QoS) under such conditions. Deep Reinforcement Learning (DRL) offers an adaptive decision-making
framework capable of learning dynamic resource allocation strategies directly from complex environments. This paper
investigates adaptive edge resource management through DRL-driven optimization models for computation offloading, task
scheduling, bandwidth allocation, energy efficiency, and autonomous orchestration in distributed edge ecosystems. The
study synthesizes recent advances between 2020 and 2025 across edge intelligence, federated learning, multi-agent
reinforcement learning, and AI-driven autonomous networking. A layered DRL-enabled edge orchestration framework is
proposed to optimize latency, throughput, energy consumption, and load balancing simultaneously. The research also
formulates two research questions focused on scalability and adaptive scheduling under heterogeneous workloads. The
proposed methodology integrates Proximal Policy Optimization (PPO), Deep Q-Networks (DQN), Multi-Agent Deep
Deterministic Policy Gradient (MADDPG), and federated reinforcement learning within a cloud-edge continuum.
Comparative analysis indicates that DRL-based adaptive management substantially improves response latency, energy
utilization, and computational efficiency compared with static and rule-based schedulers. The paper identifies unresolved
challenges involving reward engineering, explainability, convergence stability, privacy preservation, and large-scale
deployment in 6G-enabled edge systems. The findings demonstrate that DRL-driven adaptive orchestration can become a
central mechanism for autonomous edge intelligence in next-generation AI-native communication infrastructures.
This work proposes a scalable, intelligent, and resilient foundation for next-generation high-performance analytics and data-intensive applications that integrates adaptive resource management, intelligent workload scheduling, dynamic task migration, predictive analytics, and machine learning-based optimization to improve computational efficiency and responsiveness.
John Peterson, L. Martínez· International Journal of App...· 0 citations
Reinforcement learning-based adaptive resource management framework is proposed that enables cloud systems to autonomously learn optimal resource allocation policies through continuous interaction with the environment and significantly outperforms static and reactive baseline strategies in terms of resource utilization efficiency and response time stability.
Rajesh Sharma, Priya Natarajan· International Journal of Mac...· 0 citations
Modern large-scale data pipelines support analytics, AI, ML, and real-time applications but face challenges related to scalability, resource utilization, reliability, and changing workloads. This paper proposes a reinforcement learning (RL)-based autonomous optimization framework that integrates RL agents with data orchestration platforms to continuously monitor pipeline states and optimize operations. The framework uses system metrics such as workload patterns, queue lengths, execution delays, resource consumption, and failure rates to make intelligent decisions on task scheduling, resource allocation, workload balancing, and fault recovery. Three RL algorithms—Q-learning, Deep Q-Networks (DQN), and Proximal Policy Optimization (PPO)—are evaluated. Experimental results demonstrate improved throughput, reduced latency, enhanced fault tolerance, and better resource efficiency compared to traditional rule-based approaches. The proposed framework enables adaptive, self-managing data pipelines that improve scalability, resilience, and operational efficiency across enterprise, cloud, and edge environments.
Rahul Mehta· International Journal of App...· 0 citations
The research findings indicate that the key to enhancing real-time cloud network intelligence lies in the architecture based on Deep Reinforcement Learning (DRL), and provide useful guidelines for future engineering projects to ensure that network management infrastructure possesses autonomy, flexibility, and efficiency.
Shuyao Jia· The 2026 International Confe...· 0 citations
Comparative tests with PPO, FIFO, FAIR and HAS baselines confirm that multi-agent reinforcement learning can well capture the intrinsic scheduling patterns of complex mobile environments, providing an adaptive and energy-efficient scheduling solution for practical IoT deployments.
Haoyu Gu· Scientific Journal of Intell...· 0 citations