Skip to content

Advanced deep reinforcement learning techniques for dynamic resource allocation in 5G heterogeneous networks

Jul 2026 · Wireless networks · Vol 32, pp. 2375 - 2400 · 0 citations · 43 references

TL;DR

The proposed modified Deep Reinforcement Learning-based intelligent TDD configuration framework for adaptive radio resource allocation in 5G HetNets effectively enhances network reliability, resource utilization, and communication efficiency in dynamic 5G HetNet environments.

View source

Similar papers

Open access Aug 2026

Hybrid Deep Reinforcement Learning with Kookaburra Optimization for QOS-Aware Bandwidth Allocation in GMPLS Optical Networks

Objectives: To address dynamic bandwidth allocation with strict Quality of Service (QoS) requirements in Generalized Multi-Protocol Label Switching (GMPLS) optical networks under strain from internet services, real-time multimedia, and cloud infrastructure. Method: A Hybrid Deep Reinforcement Learning (Hyb-DRL) framework combined with the Kookaburra Optimization Algorithm (KkOA) for adaptive weight adjustment is proposed. Dynamically generated input data, including user request rates, queue lengths, server availability, and link stability metrics, were used to simulate real-world traffic. The Hyb-DRL agent learned optimal routing and bandwidth provisioning policies while KkOA optimized model weights for faster convergence and stability. Findings: The simulation results show that the suggested Hyb-DRL-KkOA algorithm performs better than the conventional bandwidth allocation algorithms. In contrast to conventional algorithms, it reduces the blocking probability, makespan, cost, and energy utilization while improving the throughput; therefore, providing enhanced quality of service (QoS). The proposed framework achieves a lower blocking probability by 78%, makespan by 64%, energy consumption by 51%, and operational cost by 47%. In addition to this, it provides better throughput performance by 69% than other conventional techniques. Moreover, it provided a delay of 0.0189 s, minimal energy consumption of 33 mJ, and maximal throughput of 950 Mbps. Novelty: A combination of reinforcement learning and meta-heuristic optimization leads to adaptive decision-making regarding routing and bandwidth allocation in the face of different traffic demands. The performance gain in terms of QoS is due to optimal utilization of network resources with low blocking probability, energy and operational cost. It provides a scalable and adaptive solution for high-speed, reliable data transmission in modern communication networks. Keywords: GMPLS Optical Networks, Kookaburra Optimization Algorithm, Bandwidth Allocation, Quality of Service (QoS), Blocking Probability

M. Rajagopal, S. Malathi · 0 citations
Conference Jul 2026

Deep Reinforcement Learning for Dynamic Spectrum Allocation in Cognitive Radio Network

The wildest boom of wireless devices and the shift to 5G/6G ecosystems contributed to the lack of the spectrum, making the old traditional methods of static allocation less and less efficient. cognitive radio networks provide an alternative with dynamic nature, the current solutions tend to fail because of the sophisticated nature of imperfect channel state information, large-dimensional state space and the rapid mobility of users. This model presents an Attention-Augmented Multi-agent Deep Reinforcement Learning model, which is used to maximize autonomous spectrum sharing by using spatial-temporal awareness. The architecture is based on convolutional neural networks in mapping spatial interference and long short-term memory layers in temporal mobility tracking with a multi-head self-attention mechanism to coordinate interference management between secondary users. To obtain accurate resource mapping, layers of Sinkhorn are incorporated to be bi-stochastic. Simulation shows that the spectral efficiency is 22.14% higher and the collision rate is also 6.82 times lower with the 3GPP channel models than with regular deep Q-networks. The system can be 85.36% efficient even in the presence of serious channel state errors, which is a strong solution to ensure trustworthy ultra-dense urban connectivity.

J. P. Dharshini, L. Subi, Lydia D. Isaac et al. · 0 citations
Open access Aug 2026

DEEP REINFORCEMENT LEARNING-BASED DYNAMIC SPECTRUM ACCESS FOR 6G HETEROGENEOUS COGNITIVE RADIO NETWORKS

The proliferation of heterogeneous radio access technologies in sixth-generation (6G) wireless networks demands a fundamental rethinking of spectrum management strategies. Traditional spectrum sensing approaches, designed for relatively static channel conditions, are inadequate for the dynamic, interference-rich environments that characterize 6G deployments spanning sub-6 GHz, millimetre-wave, and terahertz bands simultaneously. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users (SUs) learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models. Specifically, a Double Deep Q-Network (DDQN) architecture is adopted, augmented with a prioritized experience replay mechanism that accelerates policy convergence under non-stationary channel conditions. The proposed agent observes a composite state space encoding instantaneous channel occupancy, signal-to-interference-plus-noise ratio (SINR), primary user (PU) activity patterns, and residual energy levels, and selects actions that jointly optimize spectrum utilization efficiency, interference avoidance, and energy consumption. Simulation experiments conducted over a heterogeneous network topology with four primary users and eight secondary users demonstrate that the proposed DDQN-based scheme achieves a throughput gain of approximately 34% over conventional energy detection-based sensing, reduces interference to primary users by 61%, and attains a detection probability of 0.94 at a false alarm rate of 0.05. These results confirm the practical viability of DRL as a spectrum management backbone for next-generation cognitive radio systems.

Naadir Kamal, R. Kumar · 0 citations
Open access Jul 2026

A unified machine learning framework for intelligent resource allocation toward 6G wireless communications.

A Dual-Stage Multi-Time-Scale Temporal Attention-Based LSTM network (D-MTSTA-LSTM) has been architected, which effectively learns short- and long-term relationships in network trends, thereby precisely predicting optimal communication routes and associated power and spectrum allocation.

Nishu Gupta, Rupali Bhartiya, S. Rathod et al. · 0 citations
Preprint Aug 2026

Deep Reinforcement Learning Orchestration of Game-Theoretic User Association and Resource Allocation in HetNets

Managing dynamic User Association and Resource Allocation (UARA) in modern Heterogeneous Cellular Networks (HetNets) remains a critical open challenge. Existing mathematical optimization and Reinforcement Learning approaches face limitations in handling low-latency decision-making under dynamic traffic conditions. This paper introduces a novel orchestration scheme for game-theoretic UARA in HetNets. The proposed bilevel framework distributes UARA decisions to User Equipment through a multi-objective non-cooperative game. Overlaying the distributed game, a centralized Deep Reinforcement Learning controller orchestrates network performance by dynamically configuring the game's utility parameters, enabling transitions between power awareness, coverage enhancement, and balanced operation. Evaluated on urban HetNet topologies with 3GPP TR 38.901-compliant channel modeling, the proposed framework closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods. Furthermore, it incurs low computational overhead and maintains stable performance across the evaluated traffic densities without retraining.

Sotiris Kopsinos, Alexandros I. Papadopoulos, Antonios Lalas et al. · 0 citations
Open access Aug 2026

Dynamic Path Selection in SDN Based on Reinforcement Learning and Link Utilization

A path selection model that combines bottleneck link usage and reinforcement learning that achieves superior state awareness and adaptive routing performance in multi-source heterogeneous networks and hence can be used effectively for intelligent routing in next-generation power communication networks.

Ying Zeng, Xingnan Li, Yubeng Bao et al. · 0 citations