Skip to content
Conference

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Jul 2026 · 2026 8th International Conference on Electronics and Communication, Network and Computer Technology (ECNCT) · pp. 355-362 · 0 citations · 16 references

Abstract

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

View source

Similar papers

Open access Jul 2026

Dynamic Uplink Power Control for Cell-Free Massive MIMO

The proposed framework does not optimize only computational speed, but also clarifies the trade-off among execution time, SINR, spectral efficiency, and fairness under dynamic uplink CF-mMIMO conditions, indicating that this architecture serves as an adaptable platform to evaluate dynamic uplink power distribution across CF-mMIMO networks.

Hussein A. Jasim, M. F. A. Rasid, F. Hashim et al. · 0 citations
2026

Domain-Guided Soft Actor–Critic for Network Slicing in Cell-Free Massive MIMO Systems

Cell-free massive multiple-input multiple-output (mMIMO), which eliminates cell edge effects and enhances coverage and resource utilization, is suited for industrial Internet of things (IIoT) applications. In user-centric cell-free mMIMO-based IIoT networks, joint optimization of network slicing and access point (AP) selection is crucial for meeting diverse quality-of-service (QoS) requirements. However, the joint optimization is challenging due to the coupling of resource allocation decisions and typically imperfect channel state information. In this paper, we formulate the joint AP selection and network slicing problem as a constrained Markov decision process (CMDP) with a hybrid action space, and propose a deep reinforcement learning (RL) algorithm, domain-guided hybrid soft actor-critic for CMDP (DG-HSA2C), to maximize the long-term proportional fairness in UE transmission rates while ensuring their QoS across slices. DG-HSA2C integrates CMDP-based RL into a hybrid action space by extending the Lagrangian multiplier method. To mitigate reward hacking, our algorithm incrementally predicts future states and incorporates a domain-adaptation mechanism, enhancing fairness in resource allocation and balancing performance across slices. Simulations verify our algorithm’s effectiveness in achieving rate fairness among UEs and mitigating reward hacking under the balance of QoS and rewards.

Na Li, Meiyan Song, Hangguan Shan et al. · 0 citations
2026

Structured Reinforcement Learning for User Admission in Multi-Cell Massive MIMO via O-RAN

Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.

Jinho Choi · 0 citations
Open access Aug 2026

Energy-Efficient Cooperative Data Offloading in Cellular Networks Using Reinforcement Learning

This research is among the first to employ MARL to this extent, and it offers an end-to-end solution that combines cellular, Wi-Fi, and device-to-device (D2D) communications and considers practical network environments like user mobility and channel conditions.

Nabeel Abdolrazagh Yaseen Alrashedi, Rasool Sadeghi, Wael Hussein Zayer Al-Lamy et al. · 0 citations
Open access Jul 2026

Energy-Efficient Power Allocation for Cell-Free Massive MIMO Systems

Validation of the APG algorithm's resilience revealed that it outperformed benchmark algorithms in terms of energy efficiency and execution time, demonstrating its usefulness for challenging optimization tasks, particularly those involving bursty communication.

K. A. Bonsu, Ebenezer Baidoo Baidoo Bediako, K. Darkwah et al. · 0 citations