Skip to content
Review Open access

Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review

Jun 2026 · Future Internet · Vol 18, pp. 340 · 0 citations · 145 references

TL;DR

According to this study, PPO provides continuous action spaces with good training stability for AI models and its stable policy-learning capabilities make it suitable for next-generation communication systems.

Abstract

Fifth-generation (5G), Beyond 5G (B5G), and sixth-generation (6G) wireless networks, along with the Internet of Things (IoT), are core communication infrastructure in smart cities. Their increased deployments create high-dimensional optimization and resource management challenges. Consequently, researchers have increasingly explored the use of Artificial Intelligence (AI) models for optimizing networks. The Proximal Policy Optimization (PPO) is one such algorithm that optimizes networks. This Systematic Literature Review (SLR) follows the PRISMA 2020 protocol to review 76 studies published between 2023 and 2026 to synthesize recent PPO-based approaches to optimize communication systems. This study examines key PPO variants in major communication domains. It outlines the primary obstacles to real-world deployment and provides a cross-domain classification. According to this study, PPO provides continuous action spaces with good training stability for AI models. Its stable policy-learning capabilities make it suitable for next-generation communication systems. However, sim-to-real transfer, reward design, and multi-agent scalability are a few key challenges encountered. Future directions emphasize robust, deployable PPO frameworks for 6G, IoT, and internet architecture.

Read PDF

Similar papers

Review Open access Jul 2026

Towards Intelligent 6G Networks: A Comprehensive Review of AI-Driven Control and Optimization

The findings indicate that while AI techniques substantially improve network adaptability, resource management, and autonomous operation, significant challenges remain regarding scalability, computational complexity, data dependency, interoperability, explainability, and deployment in real-world environments.

Ali Ahmed Mirza, T. Mahmood, E. Dhulkefl · 0 citations
Conference Jul 2026

Learning-Based Resource Allocation in 5G NR Mode-2 Sidelink for Industrial AGV and AMR Communications

As the manufacturing sector increasingly adopts Industry 4.0 technologies, the need for reliable communication among devices, such as autonomous mobile robots (AMRs), or automated guided vehicles (AGVs), becomes a fundamental necessity. To support direct device-to-device communication, the third-generation partnership project (3GPP) introduced 5G new radio sidelink communication mode 2 (NR-SL) in Release 16, and 17. NR-SL allows devices to select transmission resources based on local channel sensing. In NR-SL, however, autonomous resource selection can lead to collisions, particularly in dense industrial environments. In this paper, we propose a learningbased resource allocation scheme (LBRA) for 5G NR Mode-2 sidelink. LBRA is designed to support industrial AGV and AMR communications. It employs multi-agent reinforcement learning (MARL) to improve resource allocation in NR-SL and reduces the collision probability. The results show that LBRA decreases the collision probability by approximately 73% compared to NR-SL.

Mahmoud Elsharief, Kiwoong Park, Han-Shin Jo · 0 citations
Conference Jul 2026

WHO: World Model Approach for Handover Optimization in 5G Networks

Recently, the development and deployment of intelligent controllers for radio access networks (RAN) has attracted significant attention from network operators and international telecommunications organizations, driven by rapid advances in artificial intelligence. Mobility management plays a fundamental role in ensuring seamless connectivity and service quality in 5G RAN. In fact, optimal control in 5G RAN is highly challenging due to its complex, dynamic, and distributed environment. Many approaches have been proposed to address this problem, particularly those based on deep reinforcement learning (DRL). However, contrary to the dense reward assumption in many DRL-based studies, mobility feedback in practical RAN environments is characteristically sparse and delayed. In this paper, we propose WHO (World Model for Handover Optimization), a novel method designed to bridge the gap between sparse feedback and efficient learning in 5G networks. WHO utilizes a world model to convert event-driven rewards into dense predictive signals, facilitating robust multi-agent optimization. Field experiments involving 13 base stations and 39 cells show that the proposed method significantly improves handover performance and network stability compared to conventional DRL approaches, achieving 19–40% higher prediction precision and up to 32% improvement in key performance indicators (KPIs).

Uyen Thi Thu Truong, Doan Van Nguyen, Do Ngoc Tuan et al. · 0 citations
Review Open access Jul 2026

Evolution of Power Allocation Techniques in NOMA: Advancing 5G Toward 6G Networks

The transition of 5G and beyond wireless networks toward intelligence-driven and autonomous operation has revitalized strong interest in Non-Orthogonal Multiple Access (NOMA) as an efficient multiple access framework. Power allocation critically governs NOMA performance, directly impacting throughput, user fairness, and SIC effectiveness. This survey presents a focused review of power allocation strategies in NOMA, with emphasis on the progression from static and optimization-based dynamic schemes to data-driven Artificial Intelligence (AI) and Machine Learning (ML) driven approaches. In contrast to conventional strategies that require instantaneous channel state information and iterative optimization, AI/ML techniques enable adaptive, scalable, and low-latency decision-making in highly dynamic and nonconvex environments. Recent advances in reinforcement learning and deep learning for NOMA power control are discussed, highlighting key challenges like imperfect CSI, inter-cluster interference, and distributed learning constraints. This survey provides a concise AI-centric analysis and identifies promising directions for a practical learning-driven NOMA power allocation framework for future wireless networks. A consolidated, critically comparative analysis of NOMA power allocation that bridges the gap between 5G practice and 6G imperatives is also presented in this survey.

Lekshmi Nair M, Neelakantan Pc · 0 citations
Conference Jul 2026

LLM-Assisted Network Management for Digital Twin-Enabled 5G Systems

Efficient long-term network evolution is becoming increasingly critical in dense 5G-Advanced and beyond cellular systems, where persistent traffic imbalances and localized congestion pose significant challenges that conventional short-term radio resource management alone cannot fully mitigate. This paper proposes a digital twin (DT)-enabled non-real-time (NRT) network evolution framework integrated with a large language model (LLM). Within this architecture, the digital twin provides a high-fidelity, controllable environment for evaluating infrastructure actions, while the LLM serves as a strategic orchestration engine that recommends cost-efficient network upgrades based on observed network states. Unlike traditional optimization methods that require exhaustive mathematical reformulations for each specific scenario, the proposed framework leverages the reasoning capabilities of LLMs to interpret operator objectives and constraints in natural language, generating structured evolution plans. The considered NRT action space encompasses antenna upgrades, bandwidth expansion, and new base station (BS) deployment. A techno-economic formulation is introduced to jointly evaluate load reduction performance and overall economic expenditure. Numerical results in a dense cellular scenario demonstrate that the framework effectively reduces peak resource utilization and provides diverse, coordinated evolution strategies tailored to varying network conditions.

Yukai Wang, Janghee Woo, G. Hahm et al. · 0 citations