Skip to content
Open access

Reinforcement Learning-Based Adaptive Power and Resource Allocation in Wireless Communication Networks

2026 · International Journal Of Engineering And Computer Science · 0 citations

TL;DR

This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.

Abstract

The densification of wireless networks and growing real-time service demands have intensified the need for intelligent, energy-efficient resource allocation. Traditional static and centralized methods fall short in adapting to the dynamic and interference-prone nature of 5G and emerging 6G environments. This study proposes a decentralized reinforcement learning (RL)-based framework for joint power and spectrum allocation in ultra-dense wireless systems. Each base station acts as an autonomous agent, making real-time decisions based on local traffic and interference conditions. Simulated using a custom Python-based environment with 50 base stations and 500 users, the RL approach is benchmarked against static and optimization-based methods. Results show the RL model achieves up to 91% energy efficiency, 94% spectrum utilization, and only 5% QoS degradation, outperforming baseline models. This work demonstrates the viability of RL for distributed resource management and provides a reproducible simulation toolkit to support further research in AI-driven wireless communication systems.

Read PDF

Similar papers

Open access Jul 2026

A unified machine learning framework for intelligent resource allocation toward 6G wireless communications.

A Dual-Stage Multi-Time-Scale Temporal Attention-Based LSTM network (D-MTSTA-LSTM) has been architected, which effectively learns short- and long-term relationships in network trends, thereby precisely predicting optimal communication routes and associated power and spectrum allocation.

Nishu Gupta, Rupali Bhartiya, S. Rathod et al. · 0 citations
Jul 2026

Deep Reinforcement Learning for Autonomous Communication Networks: Resource Allocation, Spectrum Management, and Control

ABSTRACT Autonomous communication systems are evolving toward self-organizing, adaptive networks capable of optimizing performance under dynamic and uncertain environments. Traditional rule-based and model-driven optimization techniques struggle to cope with the complexity, scale, and non-stationarity of modern wireless and networked systems. Reinforcement learning (RL), a branch of machine learning where agents learn optimal policies through interaction with the environment, has emerged as a powerful paradigm for enabling autonomy in communication systems. This paper (or study) explores the application of reinforcement learning techniques to autonomous communication networks, including resource allocation, spectrum management, power control, routing, and congestion control. By formulating communication tasks as Markov Decision Processes (MDPs), RL agents can learn to maximize long-term performance metrics such as throughput, latency, energy efficiency, and quality of service without requiring explicit mathematical models of the environment. Deep reinforcement learning (DRL), which integrates deep neural networks with RL, further enhances scalability by handling high-dimensional state and action spaces typical in modern networks such as 5G, 6G, and Internet of Things (IoT) systems. Multi-agent reinforcement learning (MARL) is also increasingly relevant, enabling distributed decision-making among multiple network nodes with partial observability and limited coordination. Despite its promise, RL-based communication systems face challenges including sample inefficiency, convergence stability, safety constraints, and real-time deployment limitations. Ongoing research focuses on improving training efficiency, incorporating domain knowledge, ensuring reliability, and developing hybrid models that combine RL with optimization and control theory. Overall, reinforcement learning provides a foundational framework for next-generation autonomous communication systems, enabling adaptive, intelligent, and self-optimizing networks. Keywords: Reinforcement Learning, Autonomous Communication Systems, Deep Reinforcement Learning, Multi-Agent Systems, Wireless Networks, Resource Allocation, Spectrum Management, Markov Decision Process, 5G/6G Networks, Internet of Things (IoT), Network Optimization, Self-Organizing Networks, Policy Learning, Dynamic Systems Optimization

D. A. Kumar, Jakkula Rakshitha, Madugula Pranush · 0 citations
Open access Aug 2026

Energy-Efficient Cooperative Data Offloading in Cellular Networks Using Reinforcement Learning

This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.

Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al. · 0 citations
Conference Aug 2026

Quantum Reinforcement Learning Driven Adaptive Resource Allocation for Internet of Things Devices

The high rate of Internet of Things (IoT) networks development has posed a serious problem of effective resource allocation because devices are heterogeneous, the traffic conditions are dynamic, and the energy and latency requirements are severe. Traditional resource allocation methods and classical reinforcement learning methods are not always the best methods to perform in a highly dynamic environment because they lack the adaptability and reduce convergence. The proposed paper introduces a Quantum Reinforcement Learning (QRL)-motivated adaptive resource allocation model which uses quantum-inspired state representation, as well as quantum-classical policy optimization, to optimize resource allocation. The proposed model is a dynamic distribution of bandwidth, transmission power, and computational resources the real-time state of network. Experimental assessment shows better performance than the traditional and classical reinforcement learning techniques. The proposed QRL framework attains the average accuracy of resource allocation 97.84%, lessens the communication latency by 43%, escalates the throughput by 64%, and lessens the energy usage by 37%. The system also has better convergence and stability exception when applied to different network load. Combination of quantum feature encoding improves efficacy in decision and learning. The findings affirm that the suggested framework offers a powerful and scalable system of smart resource management in the next-generation IoT systems.

A.Mohan Kumar, M. Al-Shalout, M. Elakiya et al. · 0 citations
Conference Jul 2026

Context-Aware Reinforcement Hyper-Heuristic Allocation for Dynamic Wireless Resource Management

Dynamic wireless resource allocation in multi-cell networks is challenging due to non-stationary traffic, intercell interference coupling, and heterogeneous quality-of-service (QoS) constraints. Conventional schedulers and standalone metaheuristics lack adaptability across operating regimes, while deep reinforcement learning (DRL) methods often incur high training complexity and stability limitations. This paper proposes a context-aware reinforcement hyper-heuristic framework for dynamic wireless resource allocation. A contextual bandit controller hierarchically selects among multiple low-level optimization heuristics based on real-time network state features. A multi-objective reward design jointly optimizes throughput, fairness, power efficiency, and allocation stability. We establish sublinear regret guarantees under the contextual bandit model and prove convergence under standard stochastic approximation conditions. Extensive simulations over 5,000 large-scale multi-cell instances demonstrate consistent improvements over proportional fair scheduling, evolutionary methods, and DRL-based allocators in throughput, Jain's fairness index, convergence speed, and robustness to traffic perturbations. Statistical tests confirm the significance of the gains. The results indicate that reinforcementdriven hyper-heuristic orchestration provides a scalable and theoretically grounded solution for dynamic wireless resource management.

K. Danach, Samir Haddad, J. Sayah et al. · 0 citations
Review Open access Jul 2026

Evolution of Power Allocation Techniques in NOMA: Advancing 5G Toward 6G Networks

The transition of 5G and beyond wireless networks toward intelligence-driven and autonomous operation has revitalized strong interest in Non-Orthogonal Multiple Access (NOMA) as an efficient multiple access framework. Power allocation critically governs NOMA performance, directly impacting throughput, user fairness, and SIC effectiveness. This survey presents a focused review of power allocation strategies in NOMA, with emphasis on the progression from static and optimization-based dynamic schemes to data-driven Artificial Intelligence (AI) and Machine Learning (ML) driven approaches. In contrast to conventional strategies that require instantaneous channel state information and iterative optimization, AI/ML techniques enable adaptive, scalable, and low-latency decision-making in highly dynamic and nonconvex environments. Recent advances in reinforcement learning and deep learning for NOMA power control are discussed, highlighting key challenges like imperfect CSI, inter-cluster interference, and distributed learning constraints. This survey provides a concise AI-centric analysis and identifies promising directions for a practical learning-driven NOMA power allocation framework for future wireless networks. A consolidated, critically comparative analysis of NOMA power allocation that bridges the gap between 5G practice and 6G imperatives is also presented in this survey.

Lekshmi Nair M, Neelakantan Pc · 0 citations