Skip to content
Open access

An Actor-Critic Deep Reinforcement Learning Model with Energy-Awareness and Latency Minimization for Dynamic Spectrum Allocation in 6G-Enabled Aerial Mobile Wireless Networks

Jun 2026 · Journal of Soft Computing and Data Mining · Vol 7 · 0 citations

TL;DR

An Actor-Critic Deep Reinforcement Learning (AC-DRL) model adapted to a swarm-based behavior model for dynamic spectrum allocation with energy awareness and latency minimization in AMWN is presented.

Abstract

Aerial Mobile Wireless Networks (AMWN) play an important role in next-generation communication networks, especially in military operations, disaster recovery, and real-time surveillance. However, the highly dynamic nature of AMWN creates significant challenges for dynamic spectrum allocation(DSA), latency control, energy conservation, and spectrum optimization. These challenges become more critical in 6G-enabled environments that require ultra-low latency, high energy efficiency, and large-scale device connectivity. Deep Reinforcement Learning (DRL) offers a powerful approach by enabling real-time, data-driven decision-making in complex environments. This paper presents an Actor-Critic Deep Reinforcement Learning (AC-DRL) model adapted to a swarm-based behavior model for dynamic spectrum allocationwith energy awareness and latency minimization in AMWN. The AC-DRL model is evaluated against a standard Q-learning approach using a custom dataset. The results show improved latency reduction, spectral efficiency, and energy consumption. Simulation results demonstrate up to 27% latency reduction and 22% improvement in energy efficiency compared with traditional models.

Read PDF

Similar papers

Conference Jul 2026

Multi-State Parameter Self-Optimization Using Deep Reinforcement Learning for Energy-Efficient Low-Latency Massive MIMO Systems

Massive MIMO systems require simultaneous optimization of energy efficiency, latency, and handover performance, yet existing approaches address these objectives in isolation across disparate parameter spaces. This paper proposes a multi-state parameter self-optimization framework that jointly optimizes across five interdependent operational states—channel, mobility, system configuration, power, and latency—using deep reinforcement learning. We formulate the problem as a multi-objective Markov decision process and implement five optimization approaches: Hybrid Action Space Reinforcement Learning, Q-Learning with Kalman Filter prediction, LSTM Autoencoder for PAPR reduction, bio-inspired Integrated Fruit Fly Salp Swarm Optimization for power allocation, and a proposed Multi-Agent Deep Q-Network (MA-DQN) with experience replay. Simulation results across antenna configurations from 16 to 256 elements and user counts from 5 to 40 show that the proposed MA-DQN achieves a composite performance score of $83 \pm 1.8 / 100$ across all five states (averaged over 10 seeded runs), outperforming the best single-objective method by $\mathbf{2 6} \boldsymbol{\%}$. The framework delivers 29-73% energy efficiency improvement over fixed baselines, with the learned policy favoring moderate power (0.1-0.5W) and lower antenna counts (16-32)—consistent with analytical models that show circuit power dominance at high antenna counts.

Madhu Kumari Ray, Sasmita Mohapatra, C. J. · 0 citations
Jul 2026

Deep Reinforcement Learning for Autonomous Communication Networks: Resource Allocation, Spectrum Management, and Control

ABSTRACT Autonomous communication systems are evolving toward self-organizing, adaptive networks capable of optimizing performance under dynamic and uncertain environments. Traditional rule-based and model-driven optimization techniques struggle to cope with the complexity, scale, and non-stationarity of modern wireless and networked systems. Reinforcement learning (RL), a branch of machine learning where agents learn optimal policies through interaction with the environment, has emerged as a powerful paradigm for enabling autonomy in communication systems. This paper (or study) explores the application of reinforcement learning techniques to autonomous communication networks, including resource allocation, spectrum management, power control, routing, and congestion control. By formulating communication tasks as Markov Decision Processes (MDPs), RL agents can learn to maximize long-term performance metrics such as throughput, latency, energy efficiency, and quality of service without requiring explicit mathematical models of the environment. Deep reinforcement learning (DRL), which integrates deep neural networks with RL, further enhances scalability by handling high-dimensional state and action spaces typical in modern networks such as 5G, 6G, and Internet of Things (IoT) systems. Multi-agent reinforcement learning (MARL) is also increasingly relevant, enabling distributed decision-making among multiple network nodes with partial observability and limited coordination. Despite its promise, RL-based communication systems face challenges including sample inefficiency, convergence stability, safety constraints, and real-time deployment limitations. Ongoing research focuses on improving training efficiency, incorporating domain knowledge, ensuring reliability, and developing hybrid models that combine RL with optimization and control theory. Overall, reinforcement learning provides a foundational framework for next-generation autonomous communication systems, enabling adaptive, intelligent, and self-optimizing networks. Keywords: Reinforcement Learning, Autonomous Communication Systems, Deep Reinforcement Learning, Multi-Agent Systems, Wireless Networks, Resource Allocation, Spectrum Management, Markov Decision Process, 5G/6G Networks, Internet of Things (IoT), Network Optimization, Self-Organizing Networks, Policy Learning, Dynamic Systems Optimization

D. A. Kumar, Jakkula Rakshitha, Madugula Pranush · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access Aug 2026

DEEP REINFORCEMENT LEARNING-BASED DYNAMIC SPECTRUM ACCESS FOR 6G HETEROGENEOUS COGNITIVE RADIO NETWORKS

The proliferation of heterogeneous radio access technologies in sixth-generation (6G) wireless networks demands a fundamental rethinking of spectrum management strategies. Traditional spectrum sensing approaches, designed for relatively static channel conditions, are inadequate for the dynamic, interference-rich environments that characterize 6G deployments spanning sub-6 GHz, millimetre-wave, and terahertz bands simultaneously. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users (SUs) learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models. Specifically, a Double Deep Q-Network (DDQN) architecture is adopted, augmented with a prioritized experience replay mechanism that accelerates policy convergence under non-stationary channel conditions. The proposed agent observes a composite state space encoding instantaneous channel occupancy, signal-to-interference-plus-noise ratio (SINR), primary user (PU) activity patterns, and residual energy levels, and selects actions that jointly optimize spectrum utilization efficiency, interference avoidance, and energy consumption. Simulation experiments conducted over a heterogeneous network topology with four primary users and eight secondary users demonstrate that the proposed DDQN-based scheme achieves a throughput gain of approximately 34% over conventional energy detection-based sensing, reduces interference to primary users by 61%, and attains a detection probability of 0.94 at a false alarm rate of 0.05. These results confirm the practical viability of DRL as a spectrum management backbone for next-generation cognitive radio systems.

Naadir Kamal, R. Kumar · 0 citations
Jul 2026

Advanced deep reinforcement learning techniques for dynamic resource allocation in 5G heterogeneous networks

The proposed modified Deep Reinforcement Learning-based intelligent TDD configuration framework for adaptive radio resource allocation in 5G HetNets effectively enhances network reliability, resource utilization, and communication efficiency in dynamic 5G HetNet environments.

G. Dalton, ·. A. Bamila, Virgin Louis et al. · 0 citations
Open access Jul 2026

A dynamic reward framework for scalable and efficient IoT-WSN routing using deep reinforcement learning.

A dynamic reward structuring framework within deep reinforcement learning to enable adaptive and balanced routing in IoT-WSNs and achieves significant performance gains, including approximately 30% improvement in energy efficiency, 25% reduction in latency, and 35% increase in network throughput compared with baseline methods.

Suresh Betam, S. Nagendram, Bathula Prasanna Kumar et al. · 0 citations