Skip to content
Open access

Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities

Jul 2026 · IoT · Vol 7, pp. 52 · 0 citations · 59 references

TL;DR

APO-DRL (Adaptive Protocol Optimization using Deep Reinforcement Learning), a framework that utilizes a Dueling Double Deep Q-Network combined with a Prioritized Experience Replay mechanism for intelligent, real-time communication protocol selection and parameter optimization in heterogeneous IoT networks, is presented.

Abstract

The rapid proliferation of Internet of Things (IoT) devices within 5G-enabled smart city environments has introduced unprecedented challenges in communication protocol management across heterogeneous network architectures. With connected IoT devices projected to reach 21.1 billion by the end of 2025 and approximately 39 billion by 2030, existing static protocol selection mechanisms are unable to accommodate the dynamic Quality of Service (QoS) requirements of different smart city applications, such as enhanced Mobile Broadband (eMBB), Ultra-Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC). This paper presents APO-DRL (Adaptive Protocol Optimization using Deep Reinforcement Learning), a framework that utilizes a Dueling Double Deep Q-Network (D3QN) combined with a Prioritized Experience Replay mechanism for intelligent, real-time communication protocol selection and parameter optimization in heterogeneous IoT networks. The proposed framework formulates the protocol optimization problem as a Markov Decision Process (MDP), wherein the DRL agent dynamically selects the optimal communication protocol (NB-IoT, LTE-M, LTE Cat-1, or 5G NR) and adaptively tunes transmission parameters based on real-time network conditions. Experimental evaluation in a 3GPP TR 38.901 Urban Macro simulation environment with N = 30 devices demonstrates that APO-DRL achieves a 138.9% improvement in average throughput compared to Static Allocation (60.00 vs. 25.12 Mbps), while simultaneously achieving the highest QoS satisfaction (83.38%) across all methods, albeit with higher energy consumption and packet loss than Static Allocation. Relative to D3QN+PER, APO-DRL exhibits substantially lower cross-seed throughput variance (±0.88 vs. ±11.03 Mbps), confirming that QA-PER produces a more stable and reproducible learned policy.

Read PDF

Similar papers

Open access Jul 2026

RLIOT: REINFORCEMENT LEARNING - BASED NETWORK RESOURCE OPTIMIZATION USING IOT SENSOR DATA

Results confirm that reinforcement learning–based resource allocation provides a scalable and effective solution for IoT networks, particularly in environments characterized by large state spaces, dynamic network conditions, and stochastic traffic patterns.

L. Hoang, Van-Tam Hoang, Huu-Huy Ngo · 1 citation
Open access Jul 2026

Reinforcement Learning for Resource Allocation in Energy-Harvesting Cooperative IoT Networks

This work exploits the concept of cooperative communication and radio frequency-based energy-harvesting to improve the network throughput while maintaining power supply to the IoTDs and employs the reinforcement learning frameworks, particularly state–action–reward–state–action (SARSA) and Q-learning.

Olumide Alamu, T. Olwal, Emmanuel M. Migabo · 0 citations
Open access Jul 2026

A dynamic reward framework for scalable and efficient IoT-WSN routing using deep reinforcement learning.

A dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs and achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods.

Suresh Betam, S. Nagendram, Bathula Prasanna Kumar et al. · 0 citations
Review Open access Aug 2026

AI-Driven Energy-Efficient Network Slicing for UAV-Assisted 6G IoT Communications Using Deep Reinforcement Learning

Sixth-generation (6G) wireless networks are expected to support massive Internet of Things (IoT) connectivity, ultra-reliable low latency services, high-throughput multimedia traffic, and intelligent and autonomous infrastructures. Conventional terrestrial deployments may be insufficient in rural areas, disaster recovery scenarios, emergency zones, and temporary high-density IoT events, where rapid coverage extension and adaptive resource management are required. Unmanned aerial vehicles (UAVs) can operate as aerial base stations to enhance service availability; however, limited onboard energy, altitude-dependent air-to ground channels, constrained bandwidth and transmit power, and heterogeneous quality-of-service (QoS) requirements make static resource allocation inefficient. This revised paper proposes an AI driven energy-efficient network slicing framework for UAV assisted 6G IoT communication. The network is divided into enhanced Mobile Broadband (eMBB), Ultra- Reliable Low-Latency Communication (URLLC), and massive Machine-Type Communication (mMTC) slices. A DQN-based deep reinforcement learning (DRL) agent dynamically allocates the slice-level bandwidth, transmit power, and altitude-control actions after converting continuous decision variables into a finite feasible action set. The reward function jointly maximizes the throughput and energy efficiency while penalizing the latency, packet loss, and QoS violations. To address the reviewers’ concerns, the revised manuscript adds an LoS/NLoS air- to-ground channel model, a propulsion-aware UAV energy model, detailed DRL hyperparameters, a nine-action discretization table, Monte Carlo validation over 30 independent seeds, Welch significance testing, DRL variant comparison, computational complexity analysis, and three relevant references from Sana’a University Journalof Applied Sciences and Technology. The proposed method improves the throughput by 15.7%, reduces the latency by 19.8%, improves the energy efficiency by 16.4%, and reduces the packet loss by 24.6% compared with the greedy baseline. The results confirm that slice-aware DRL improves resource utilization and service reliability in UAV-assisted 6G IoT networks.

Unknown authors · 0 citations
Open access 2025

Intelligent Resource Allocation in Smart Cities Using Multi-Agent Reinforcement Learning

The rapid integration of Internet of Things (IoT), Artificial Intelligence (AI), cloud computing, edge computing, and advanced communication technologies is transforming traditional urban infrastructure into intelligent smart cities. Conventional resource allocation methods struggle to manage dynamic urban environments, creating a need for adaptive and decentralized decision-making systems. This study proposes a Multi-Agent Reinforcement Learning (MARL) framework for intelligent resource allocation across transportation, energy, water, healthcare, emergency response, and communication systems. Each urban subsystem functions as an autonomous learning agent that optimizes local decisions while coordinating to improve overall city performance. The framework combines IoT sensing, edge intelligence, cloud analytics, and deep reinforcement learning to enable real-time, adaptive resource management. It enhances resource utilization, reduces energy consumption and response time, improves system reliability, and supports scalable urban operations. The proposed approach also provides a foundation for future smart city technologies, including digital twins, federated learning, autonomous edge intelligence, and 6G networks, promoting sustainable, resilient, and efficient urban development.

Seshagiri N · 0 citations