Skip to content
Review

Deep Reinforcement Learning for 6G AI-RAN: A Comprehensive Survey

Aug 2026 · 0 citations
Computer Science Engineering

TL;DR

This article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN, and reviews the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provides an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure.

Abstract

The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous services, dynamic environments, and competing performance objectives. Open Radio Access Network (O-RAN) provides the open interfaces, disaggregated architecture, and multi-timescale control loops needed to support this transformation, while deep reinforcement learning (DRL) offers a natural framework for optimizing sequential decisions under uncertainty. However, existing surveys either address artificial intelligence (AI) and machine learning (ML) in O-RAN broadly or focus on isolated DRL use cases, leaving a gap in the systematic connection between DRL methodology, O-RAN architecture, and operational deployment. To the best of our knowledge, this article presents the first dedicated and comprehensive survey of DRL for Open AI-RAN. We review the foundations of model-free, model-based, offline, safe, multi-agent, federated, and transfer learning, and provide an O-RAN-aware framework for formulating RAN control problems through states, observations, actions, rewards, constraints, and temporal structure. We classify DRL applications across radio resource management, mobility management, interference control, traffic steering, energy efficiency, network slicing, integrated sensing and communication, security, and massive MIMO. We further examine multi-agent and federated coordination, foundation models and agentic AI, trustworthy DRL, sim-to-real transfer, continual adaptation, resource-efficient inference, and reinforcement learning operations. Finally, we review experimental platforms, benchmarks, standards, and industry activities, and identify research directions toward sample-efficient, safe, scalable, interoperable, and deployable DRL control for 6G Open AI-RAN.

View source

Similar papers

Preprint Aug 2026

EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RAN

Despite recent advances in applying artificial intelligence (AI) techniques to radio access network (RAN), critical challenges remain: traditional machine learning (ML) algorithms suffer from limited generalization across varying network topologies, whereas general-purpose large language models (LLMs) face high computational demands and lack domain-specific knowledge. To address these gaps, this article introduces the evolving RAN intelligent controller (RIC) (EvoRIC) framework, a hierarchical architecture that enables continuous evolution by leveraging a non-real-time RIC (non-RT RIC) for global model updates and a near-real-time RIC (near-RT RIC) for local execution, dynamically empowering LLMs with domain-specific decision-making capabilities. Within this framework, we employ a reinforcement learning-based fine-tuning (RLFT) mechanism where an LLM operates as an actor within a proximal policy optimization (PPO) agent. By leveraging the interaction tuples collected from the wireless environment, the LLM's parameters are iteratively updated to align semantic reasoning with rigorous network performance objectives. We evaluate the generalization and efficacy of the proposed EvoRIC framework within integrated access and backhaul (IAB) networks, and finally, discuss the open challenges and future directions of the EvoRIC framework toward realizing autonomous O-RAN.

Lingyan Bao, Jemin Lee, Tony Q. S. Quek · 0 citations
Conference Jul 2026

A Digital Twin-Based Deep Reinforcement Learning Framework for Adaptive Scheduling in 5G/6G Networks

The transition toward AI-native 6G networks requires intelligent, adaptive, and reliable control mechanisms capable of handling highly dynamic and heterogeneous environments. In this context, reinforcement learning has emerged as a promising approach for optimizing network performance. However, existing works often focus on algorithmic design while overlooking critical aspects such as experimental rigor, reward formulation, and reproducibility, which can significantly impact the validity of the results. This paper proposes a Digital Twin-based Deep Reinforcement Learning framework for adaptive scheduling in 5G networks, where the twin is implemented as a simulation-driven proxy. Within this framework, a Deep Q-Network agent dynamically selects transmission interval configurations to jointly optimize key Quality of Service metrics. A key contribution of this work lies in the systematic identification and correction of critical experimental issues, including reward degeneracy and hidden coupling between control variables, which are often neglected in RL-based networking studies. To ensure robustness, the proposed approach incorporates multi-seed statistical evaluation, providing reproducible and reliable performance assessment. Experimental results demonstrate that the proposed DRL agent consistently converges toward the optimal scheduling configuration and achieves stable performance across multiple runs, outperforming a tabular Q-learning baseline.

Naima Mchiri, Rachid Zagrouba, E. Zagrouba · 0 citations
Review Jul 2026

JEPA for AI-Native 6G: Predictive Representations and Open Challenges

Sixth-generation (6G) networks are moving toward AI-native operation, where learning modules are embedded across the radio access network (RAN), edge, and core. This transition requires learning from limited labels, heterogeneous wireless and network data, partial observations, non-stationary propagation, and latency-constrained control loops. Joint-embedding predictive architecture (JEPA) is a promising self-supervised paradigm for this setting because it predicts missing or future representations in latent space instead of reconstructing raw measurements or using contrastive negative samples. This article presents a wireless-oriented tutorial on JEPA for 6G intelligence. We define the JEPA training mechanism, describe how CSI, beam measurements, KPIs, topology graphs, and sensing observations can be tokenized and masked, and position the learned encoder as a predictive representation layer for RAN, O-RAN, edge, and core functions, with task-specific heads or controllers producing final decisions. Then we present an illustrative, beam-management case study suggesting that a wireless-aware target, specifically an auxiliary future beam-energy target during self-supervised pretraining, can improve label efficiency and robustness across shifted deployment conditions relative to a supervised source domain. Finally, we outline open challenges in multi-timescale prediction, action-conditioned modeling, distributed training, trustworthiness, efficient deployment, benchmarking, and standardization.

Sheikh Salman Hassan, Irshad A. Meer, Almoatssimbillah Saifaldawla et al. · 0 citations
Open access Aug 2026

L-ARLPT: An LLM-Augmented Reinforcement Learning Framework for Autonomous Penetration Testing

A Large Language Model-enhanced Autonomous Reinforcement Learning Penetration Testing framework that leverages the domain knowledge embedded in a Large Language Model to perform tactical planning, thereby pruning the original action space into a compact set of candidate actions.

Rufeng Zhan, Junyi Zhu, Yinghui Xu et al. · 0 citations
Open access 2026

AI-Enabled Autonomous Network Slicing Optimization for 6G Communication Systems

This paper presents a comprehensive framework for artificial intelligence (AI)-enabled autonomous network slicing optimization in 6G systems and investigates the application of advanced machine learning paradigms specifically deep reinforcement learning, federated learning, and generative AI to orchestrate dynamic resource provisioning, cross-slice isolation, and proactive SLA (Service Level Agreement) enforcement.

N. P J, Jeeva Jothi · 0 citations
Review Open access Dec 2025

Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks

Reinforcement Learning (RL) has shown remarkable success in enabling adaptive and data-driven optimization for various applications in wireless networks. However, classical RL suffers from limitations in generalization, learning feedback, interpretability, and sample efficiency in dynamic wireless environments. Large Language Models (LLMs) have emerged as a transformative Artificial Intelligence (AI) paradigm with exceptional capabilities in knowledge generalization, contextual reasoning, and interactive generation, which have demonstrated strong potential to enhance classical RL. This paper serves as a comprehensive tutorial on LLM-enhanced RL for wireless networks. We propose a taxonomy to categorize the roles of LLMs into four critical functions: state perceiver, reward designer, decision-maker, and generator. Then, we review existing studies exploring how each role of LLMs enhances different stages of the RL pipeline. Moreover, we provide a series of case studies to illustrate how to design and apply LLM-enhanced RL in low-altitude economy networking, vehicular networks, and space–air–ground integrated networks. Finally, we conclude with a discussion on potential future directions for LLM-enhanced RL and offer insights into its future development in wireless networks.

Lingyi Cai, Wenjie Fu, Yuxi Huang et al. · 2 citations