Skip to content

Domain-Guided Soft Actor–Critic for Network Slicing in Cell-Free Massive MIMO Systems

2026 · IEEE Transactions on Communications · Vol 74, pp. 11486-11503 · 0 citations · 53 references
Computer Science

Abstract

Cell-free massive multiple-input multiple-output (mMIMO), which eliminates cell edge effects and enhances coverage and resource utilization, is suited for industrial Internet of things (IIoT) applications. In user-centric cell-free mMIMO-based IIoT networks, joint optimization of network slicing and access point (AP) selection is crucial for meeting diverse quality-of-service (QoS) requirements. However, the joint optimization is challenging due to the coupling of resource allocation decisions and typically imperfect channel state information. In this paper, we formulate the joint AP selection and network slicing problem as a constrained Markov decision process (CMDP) with a hybrid action space, and propose a deep reinforcement learning (RL) algorithm, domain-guided hybrid soft actor-critic for CMDP (DG-HSA2C), to maximize the long-term proportional fairness in UE transmission rates while ensuring their QoS across slices. DG-HSA2C integrates CMDP-based RL into a hybrid action space by extending the Lagrangian multiplier method. To mitigate reward hacking, our algorithm incrementally predicts future states and incorporates a domain-adaptation mechanism, enhancing fairness in resource allocation and balancing performance across slices. Simulations verify our algorithm’s effectiveness in achieving rate fairness among UEs and mitigating reward hacking under the balance of QoS and rewards.

View source

Similar papers

Conference Jul 2026

AP Selection and Power Control for Personalized Cell-Free Massive MIMO: Graph-Embedded Reinforcement Learning Approach

Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.

Yu-Heng An · 0 citations
Conference Jul 2026

Graph-Centric Deep Q-Learning for Interference-Aware Resource Allocation in Rsma-Enabled 5G Slicing

The emergence of 5G and 6G advanced ecosystems demands highly adaptive resource management to orchestrate the specialised requirements of eMBB, URLLC, and mMTC network slices. In dense multi-cell environments, capturing complex spatial interdependencies and mitigating dynamic interference is paramount for maintaining Quality of Service (QoS). This paper introduces a robust GNN-DQN framework designed for Rate Splitting Multiple Access (RSMA) based networks. By representing the network topology as a graph, the framework leverages Graph Neural Networks (GNNs) to extract highdimensional spatial features and model inter-cell interference patterns. These insights enable a Deep Q-Network (DQN) agent to perform intelligent resource partitioning and dynamic power splitting of the RSMA common stream. Experimental results demonstrate that the proposed GNN-DQN framework achieves a connectivity success ratio exceeding 90% across all slices, representing an average improvement of over 60% compared to non-graph-based reinforcement learning and supervised baselines. Notably, the framework demonstrates exceptional spectral efficiency, maintaining near-total connectivity while utilising less than 10% of the normalised system bandwidth, a 4× reduction in resource overhead compared to traditional methods. Furthermore, the GNN-driven architecture ensures stable convergence during training, yielding a 1.6× higher system reward score. Our findings validate GNN-DQN as a high-performance, scalable, and resource-efficient paradigm for intelligent orchestration in 5G and 6G networks.

Aya Kh. Ahmed, Nadia Al-Aboody, Hamed S. Al-Raweshidy · 0 citations
2026

Structured Reinforcement Learning for User Admission in Multi-Cell Massive MIMO via O-RAN

Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.

Jinho Choi · 0 citations
Conference Jul 2026

A Unified Group-Wise SAC Framework for RIS Deployment Optimization in Cell-Free MIMO Networks

Reconfigurable intelligent surface (RIS)-aided cellfree multiple-input multiple-output (CF-MIMO) is emerging as a pivotal architecture for 6G networks, promising ubiquitous coverage and high spectral efficiency. However, the large-scale deployment of RISs entails significant challenges in network planning, specifically the joint optimization of sparse RIS deployment, phase configuration, and active beamforming. This constitutes a complex mixed-integer non-linear programming (MINLP) problem. To address this, we propose a unified deep reinforcement learning (DRL) framework, which optimizes discrete RIS selection and continuous beamforming variables in an end-to-end manner. Specifically, we introduce a Top-K relaxation mechanism to handle the binary deployment constraints and group-wise phase control strategy to mitigate the high-dimensional action space in large-scale RISs. Simulation results demonstrate that the proposed framework effectively achieves rapid convergence. Specifically, the proposed joint optimization yields a maximum sum-rate improvement of 40.4% over the network without RISs.

Bo Yin, Wout Joseph, Jorn Schampheleer et al. · 0 citations
Jul 2026

Hierarchical Multi-Objective Learning for Context-Aware 5G Ran Slice Resource Allocation

Efficient coexistence of eMBB and URLLC services remains a critical challenge in AI-native Radio Access Networks (RANs). This paper proposes a two-timescale Hierarchical Reward Weighting (HRW) framework based on multiobjective reinforcement learning for context-aware O-RAN slicing under a Constrained Markov Decision Process (CMDP) formulation. The proposed architecture separates long-term policy adaptation from fast-timescale radio scheduling, mitigating the non-stationarity inherent in multiobjective RAN optimization. At the slow layer, a non-realtime RIC rApp exploits a long-term network context and a differentiable Softmax mapping to adapt slice reward preferences. These policies are propagated through the $O$ -RAN control hierarchy to guide downstream scheduling decisions. At the fast layer, decentralized scheduling agents embedded within the Open Distributed Unit (O-DU) MAC layer execute sub-millisecond Physical Resource Block (PRB) allocation and packet preemption, avoiding near-RT RIC transport latency constraints. Evaluated under a multiuser MIMO-OFDMA environment, the proposed framework improves resource utilization by up to 60.8% over static partitioning while maintaining bounded URLLC tail-latency behavior and strict Service Level Agreement (SLA) compliance. The results demonstrate the feasibility of AI-native hierarchical O-RAN control and align with the ITU-T visions for autonomous 6G RAN intelligence.

Charles Ssengonzi, Okuthe P. Kogeda, T. Olwal · 0 citations
Open access 2026

A Profile-Aware Resource Allocation Framework for RAN Slicing in Cell-Free Massive MIMO Networks

As wireless networks transition toward the 6G era, supporting strictly heterogeneous services such as eMBB, URLLC, and mMTC over a unified infrastructure becomes a fundamental challenge. Traditional Radio Access Network (RAN) slicing often relies on upper-layer logical abstractions, which fail to address physical inter-slice interference and the boundary effect of cellular architectures. This paper proposes a novel Profile-Aware Hierarchical RAN Slicing Framework for User-Centric Cell-Free Massive MIMO systems to overcome these limitations. The proposed framework comprises four hierarchical stages: utility-based profile-aware clustering, dynamic inter-slice resource partitioning for power and bandwidth, hybrid central processing unit-to-access point power budgeting, and real-time power allocation using a Bipartite Graph Convolutional Network (BiGCN). By incorporating service-specific requirements into physical layer resource management, the framework ensures strict quality-of-service isolation and global energy efficiency. Simulation results demonstrate that the proposed integrated framework achieves a high Jain’s Fairness Index, exceeding 0.95 for most profiles and provides up to a 48-fold energy efficiency improvement for battery-constrained devices compared to non-slicing baselines. Furthermore, the BiGCN-based allocation module attains near-optimal performance with millisecond-level inference latency, confirming its feasibility for mission-critical real-time applications. This comprehensive approach effectively eliminates the trade-off between aggregate throughput and individual reliability, providing a scalable solution for next-generation sliced networks.

Ja-Eun Kim, Hye-Yoon Jeong, Ji-Woo Lee et al. · 0 citations