Jul 2026· 2026 11th International Conference on Applying New Technology in Green Buildings (ATiGB)· pp. 226-231· 0 citations· 26 references
Abstract
Massive MIMO systems require simultaneous optimization of energy efficiency, latency, and handover performance, yet existing approaches address these objectives in isolation across disparate parameter spaces. This paper proposes a multi-state parameter self-optimization framework that jointly optimizes across five interdependent operational states—channel, mobility, system configuration, power, and latency—using deep reinforcement learning. We formulate the problem as a multi-objective Markov decision process and implement five optimization approaches: Hybrid Action Space Reinforcement Learning, Q-Learning with Kalman Filter prediction, LSTM Autoencoder for PAPR reduction, bio-inspired Integrated Fruit Fly Salp Swarm Optimization for power allocation, and a proposed Multi-Agent Deep Q-Network (MA-DQN) with experience replay. Simulation results across antenna configurations from 16 to 256 elements and user counts from 5 to 40 show that the proposed MA-DQN achieves a composite performance score of $83 \pm 1.8 / 100$ across all five states (averaged over 10 seeded runs), outperforming the best single-objective method by $\mathbf{2 6} \boldsymbol{\%}$. The framework delivers 29-73% energy efficiency improvement over fixed baselines, with the learned policy favoring moderate power (0.1-0.5W) and lower antenna counts (16-32)—consistent with analytical models that show circuit power dominance at high antenna counts.
Sixth-generation (6G) mobile communication poses unprecedented challenges for resource scheduling under personalized demands. Cell-free massive multiple-input multiple-output (CF-mMIMO), with its user-centric characteristics, has emerged as a key technology for satisfying personalized demands. However, faced with heterogeneous quality-of-service (QoS) requirements, existing reinforcement learning schemes are constrained by partial observability, making it difficult to balance overall system performance and personalized demands. Consequently, we propose a graph-embedded multi-agent deep deterministic policy gradient (G-MADDPG) scheme. Guided by personalized demands, proposed G-MADDPG formulates a maximization problem for system weighted sum spectral efficiency and introduces differentiated QoS penalties. In addition, graph neural networks (GNNs) are embedded into the policy learning and value estimation processes of reinforcement learning, endowing agents with enhanced structural reception and cooperative capabilities. Simulation results demonstrate that proposed G-MADDPG scheme outperforms existing benchmark schemes in both convergence speed and performance evaluation.
Yu-Heng An· 2026 8th International Confe...· 0 citations
This work investigates the joint optimization of Age of Information (AoI) and energy harvesting (EH) in wireless edge computing systems, where edge servers not only process IoT data but also act as wireless power suppliers via simultaneous wireless information and power transfer (SWIPT). Building upon the asynchronous model-free fractional multi-agent reinforcement learning framework and the Lyapunov drift-plus-penalty (DPP) concept, we design a fractional-based reward function for AoI and construct a virtual queue to enforce long-term energy stability under battery storage constraints. The overall reward is formulated as a weighted sum, capturing the trade-off between timeliness and energy sustainability, with update decisions, task offloading, and power splitting ratios as key control variables. Simulation results demonstrate that the developed multi-agent deep reinforcement learning approach achieves superior AoI–energy trade-offs compared to related baseline algorithms. These findings highlight the effectiveness of our framework in balancing information freshness and sustainable energy harvesting under resource-constrained edge environments.
This paper proposes a system-aware adaptive channel state information (CSI) feedback framework for massive multiple-input multiple-output (mMIMO) systems, aiming to dynamically optimize the trade-off between reconstruction fidelity and signaling overhead. While deep learning-based autoencoders (AEs) have enabled significant CSI compression, conventional fixed-ratio schemes fail to adapt effectively to non-stationary channel conditions. To address this limitation, we develop a reinforcement learning (RL)-driven control framework that operates over a bank of pretrained multi-rate AEs, each corresponding to a distinct compression ratio (CR). At each time step, a centralized RL agent selects the most suitable CR for each user based on observed channel conditions and system performance indicators. Distinct from conventional mean squared error (MSE)-centric designs, we introduce a system-aware reward formulation that jointly accounts for spectral efficiency via signal-to-interference-plus-noise ratio (SINR), feedback overhead constraints, and the computational cost of model adaptation. Simulation results on high-dimensional delay-domain CSI datasets demonstrate that the proposed RL-guided framework effectively balances the overhead-accuracy tradeoff and adapts to dynamic channel environments. The proposed method improves spectral efficiency and feedback efficiency compared with fixed compression schemes and adaptive baselines, while maintaining a modest computational and memory footprint. Averaged over different numbers of users and across all considered baselines, the proposed RL framework reduces the CSI feedback cost by more than 53.4%, improves the average downlink sum rate by 53.64%, and reduces the NMSE by 22.38%. These results demonstrate its ability to achieve a more efficient rate-accuracy-feedback tradeoff under dynamic wireless conditions.
Maryam Ansarifard, M. Sharma, George Exarchakos et al.· 0 citations
The ability of various isolated devices to sense their surroundings can be improved by 5G millimetre wave (mmWave) communication technology. By jointly supporting data transmission and sensing tasks, the framework improves overall spectrum efficiency in wireless networks. Among them, the Integrated Sensing and Communication (ISAC) has become the standard in wireless communications. Specifically, mmWave technology is highly effective for bandwidth-intensive communication services and delivers improved spatial and temporal accuracy through its large spectrum availability and directional beamforming characteristics. To meet the requirements, a multi-agent-based deep learning technique is proposed for better development. Over this sensing network of 5G mmWave, the resource allocation process is handled by Multi-agent Deep Reinforcement Learning with Prioritized Experience Replay (MDRL-PER), whereas the system is provided based on allocated resource for better communication. Finally, the performance of the system is assessed through distinct evaluation metrics and compared with existing methodologies. Hence, the superior results are obtained to ensure the efficacy of the communication network.
Papisetty Sai Prasad, T. Kavitha· 2026 7th International Confe...· 0 citations
Artificial intelligence (AI) and machine learning (ML) are increasingly applied to wireless and cellular networks. With sixth-generation (6G) systems envisioned as AI-native, reinforcement learning (RL) offers a natural approach to complex network management and operation. This paper focuses on user admission control in multi-cell massive multiple-input multiple-output (MIMO) systems, where naive selfish strategies aiming to maximize local sum-rate can trigger a tragedy of the commons, degrading per-user performance and generating severe inter-cell interference (ICI). To address these challenges, we introduce a structured RL framework for massive MIMO systems. In particular, the policy is structured to introduce physical inductive bias terms, such as an interference-sensitive attenuation factor, which enables interference-aware learning through the open radio access network (O-RAN) architecture. Through stability analysis, we show that such physical inductive bias terms can guarantee network-wide stability. Experimental results demonstrate that the proposed approach balances aggregate spectral efficiency with per-user performance and maintains robustness during traffic surges, whereas selfish strategies suffer from degraded per-user performance.
Jinho Choi· IEEE Transactions on Communi...· 0 citations
ABSTRACT
Autonomous communication systems are evolving toward self-organizing, adaptive networks capable of optimizing performance under dynamic and uncertain environments. Traditional rule-based and model-driven optimization techniques struggle to cope with the complexity, scale, and non-stationarity of modern wireless and networked systems. Reinforcement learning (RL), a branch of machine learning where agents learn optimal policies through interaction with the environment, has emerged as a powerful paradigm for enabling autonomy in communication systems. This paper (or study) explores the application of reinforcement learning techniques to autonomous communication networks, including resource allocation, spectrum management, power control, routing, and congestion control. By formulating communication tasks as Markov Decision Processes (MDPs), RL agents can learn to maximize long-term performance metrics such as throughput, latency, energy efficiency, and quality of service without requiring explicit mathematical models of the environment. Deep reinforcement learning (DRL), which integrates deep neural networks with RL, further enhances scalability by handling high-dimensional state and action spaces typical in modern networks such as 5G, 6G, and Internet of Things (IoT) systems. Multi-agent reinforcement learning (MARL) is also increasingly relevant, enabling distributed decision-making among multiple network nodes with partial observability and limited coordination. Despite its promise, RL-based communication systems face challenges including sample inefficiency, convergence stability, safety constraints, and real-time deployment limitations. Ongoing research focuses on improving training efficiency, incorporating domain knowledge, ensuring reliability, and developing hybrid models that combine RL with optimization and control theory. Overall, reinforcement learning provides a foundational framework for next-generation autonomous communication systems, enabling adaptive, intelligent, and self-optimizing networks.
Keywords: Reinforcement Learning, Autonomous Communication Systems, Deep Reinforcement Learning, Multi-Agent Systems, Wireless Networks, Resource Allocation, Spectrum Management, Markov Decision Process, 5G/6G Networks, Internet of Things (IoT), Network Optimization, Self-Organizing Networks, Policy Learning, Dynamic Systems Optimization
D. A. Kumar, Jakkula Rakshitha, Madugula Pranush· International Scientific Jou...· 0 citations