Simulation results position D3QN-PER as a strong candidate for deployment as a near-RT RIC xApp within the O-RAN architecture, advancing the vision of AI-native mobility management for 6G.
This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.
Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al.· Journal of universal compute...· 0 citations
Ultra-dense 5G networks require advanced traffic steering to maintain performance and balance load amid growing user and base station (gNB) densities. Traditional heuristics such as nearest-base-station and Signal-to-Interference-plus-Noise Ratio (SINR)-based selection provide simple solutions but struggle to adapt to dynamic user mobility, diverse traffic, and fluctuating radio conditions at the mobility-control level, often leading to inefficient handovers and degraded network quality. We propose a deep reinforcement learning (DRL) framework to dynamically tune a global handover hysteresis margin that governs handover triggering decisions, optimizing handover success, reducing failures, and enhancing throughput and fairness. Implemented in Python using Stable Baselines3 and NumPy, our custom simulation environment models key mobility-related 5G dynamics at a high level, including user mobility, pathloss-based signal degradation, and interference. We evaluate DRL agents-Deep Q-Network (DQN) and Proximal Policy Optimization (PPO)-against heuristic and hysteresis-based baselines. Results show that DRL-based hysteresis optimization provides strong and robust performance under the considered ultra-dense mobility conditions in handover success rate, average SINR, throughput, and fairness, with PPO demonstrating the most consistent behavior across configurations. This work offers a reproducible simulation framework for further research into adaptive mobility management.
Damianos Diasakos, V. Kokkinos, C. Bouras et al.· International Conference on...· 0 citations
A Hybrid deep double-Q networks (DDQN)–bidirectional long short-term memory (Bi-LSTM) Framework that integrates bi-directional mobility prediction and DRL-based adaptive decision-making is introduced that highlights the effectiveness of integrating predictive intelligence with reinforcement learning for reliable mobility management in 5G-Advanced and emerging 6G networks.
Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.
Yu-Heng An· 2026 8th International Confe...· 0 citations
The proliferation of heterogeneous radio access technologies in sixth-generation (6G) wireless networks demands a fundamental rethinking of spectrum management strategies. Traditional spectrum sensing approaches, designed for relatively static channel conditions, are inadequate for the dynamic, interference-rich environments that characterize 6G deployments spanning sub-6 GHz, millimetre-wave, and terahertz bands simultaneously. This paper proposes a Deep Reinforcement Learning (DRL)-based framework for dynamic spectrum access in 6G heterogeneous Cognitive Radio Networks (Het-CRNs), wherein secondary users (SUs) learn optimal channel selection policies through direct interaction with the radio environment, without requiring explicit statistical channel models. Specifically, a Double Deep Q-Network (DDQN) architecture is adopted, augmented with a prioritized experience replay mechanism that accelerates policy convergence under non-stationary channel conditions. The proposed agent observes a composite state space encoding instantaneous channel occupancy, signal-to-interference-plus-noise ratio (SINR), primary user (PU) activity patterns, and residual energy levels, and selects actions that jointly optimize spectrum utilization efficiency, interference avoidance, and energy consumption. Simulation experiments conducted over a heterogeneous network topology with four primary users and eight secondary users demonstrate that the proposed DDQN-based scheme achieves a throughput gain of approximately 34% over conventional energy detection-based sensing, reduces interference to primary users by 61%, and attains a detection probability of 0.94 at a false alarm rate of 0.05. These results confirm the practical viability of DRL as a spectrum management backbone for next-generation cognitive radio systems.
Naadir Kamal, R. Kumar· Global Journal of Engineerin...· 0 citations
A comprehensive survey of AI-enabled mobility management strategies for 5G, Beyond 5G, and upcoming 6G networks, with particular attention to HO optimization and load balancing is presented.
H. Asif, Abdulraqeb Alhammadi, N. Tarhuni et al.· Future Internet· 0 citations