Sep 2026· IEEE Transactions on Smart Grid· Vol 17, pp. 3924-3940· 1 citation· 45 references
Abstract
The complex multi-energy coupling characteristics inherent to integrated energy system (IES) present unprecedented challenges for the implementation of low-carbon scheduling. Existing optimization methods often exhibit limitations in system scalability, algorithm adaptivity, and carbon reduction efficacy for complex IES. This paper proposes a Large Language Model (LLM)-Embedded Multi-Agent Reinforcement Learning (LEMARL) to address the aforementioned issues. The proposed method integrates the global perception capability of LLMs with the dynamic optimization capability of MARL. Specifically, the LLM-Embedded module generates high-quality reward functions and policy frameworks from a global perspective, while the MARL module leverages these LLM-generated strategies for distributed interactive iterations—greatly enhancing computation efficiency and scalability. Simulation results demonstrate that LEMARL reduces carbon emissions by 7.76% and simultaneously decreases operating costs by 4.49% in a small-scale IES. Furthermore, LEMARL also exhibits superior applicability and scalability in large-scale IES of the IEEE 141-bus power grid integrated with 51-node thermal system.
A comprehensive review of RL-based control applications across RES-integrated energy domains, including power grids, microgrids, and building energy systems, identifies emerging trends and highlights dominant design patterns across power grid, microgrid, and building-level applications.
P. Michailidis, F. Minelli, H. Coban et al.· Infrastructures· 0 citations
To enhance power grid adaptability amid rising renewable energy integration, this paper proposes M3-PPO, a meta-reinforcement learning algorithm that enables efficient the strategy transfer and rapid adaptation across tasks with varying energy mixes. Built upon a base framework (M-PPO) that integrates PPO and MAML, M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop. Experiments on the Grid2Op platform demonstrate that M3-PPO significantly outperforms baseline algorithms in generalization and scheduling efficiency, achieving robust performance even when simulating complex energy environments. The approach is particularly suitable for integration with antenna-enabled smart grid monitoring, wireless data acquisition, and edge-computing platforms, providing an engineering-oriented solution for adaptive, real-time, and robust power grid scheduling in modern renewable-rich energy systems.
Q. Dai, X. Hu, J. Li et al.· Advanced Electromagnetics· 0 citations
While traditional day-ahead economic scheduling provides a foundational baseline for microgrid energy management, the inevitable transition toward real-time, dynamic control necessitates algorithms with extreme computational efficiency. This paper presents a microgrid optimization strategy predicated on the Q-learning reinforcement learning (RL) algorithm. The multi-energy scheduling problem is formulated as a Markov Decision Process (MDP), wherein strict physical boundaries—specifically battery State of Charge (SOC) limits—are directly internalized via a tailored reward function. Operating on a decoupled "offline training and online inference" paradigm, the RL agent is evaluated under a baseline day-ahead framework to verify its global optimization capabilities. Comparative simulations against Particle Swarm Optimization (PSO) demonstrate a 27.8% enhancement in economic profitability (yielding a minimum cost of -$369.39). Crucially, the RL approach achieves an ultra-low online inference latency of merely 0.0018 s. By utilizing the day-ahead model purely as an economic benchmark, this research validates that the proposed RL paradigm not only guarantees optimal dispatch but also fundamentally shatters the computational bottlenecks of heuristic algorithms, establishing a critical algorithmic foundation for high-frequency hardware-in-the-loop simulations and multi-agent real-time coordination.
Xianzhi Chen, Yuan Li, Jiatian Zhang et al.· Advances in Engineering Tech...· 0 citations
Microgrids must efficiently manage energy under uncertainties in renewable generation and load demand to ensure reliable and cost-effective operation. This paper investigates a microgrid system that involves renewable energy through the photovoltaic system, wind system, battery energy storage, and local load requirement with a centralized energy management system. The inflexible nature of traditional rule-based and optimizationbased approaches to solving problems can often create issues in reflection to dynamic operating conditions, and reinforcement learning approaches can produce unsafe control behavior in exploration stages. To address these issues, this paper suggests a hierarchical hybrid energy management structure that will integrate rule-based supervision and a SARSA reinforcement learning controller. Supervisory layer ensures that the system is safe by ensuring that there are operational limits such as battery state of charge limits as well as power balance conditions. The learning agent on the other hand optimizes the control choices to reduce operational costs and grid energy consumption. The outcomes of the simulation indicate that the suggested approach saves more money, learns quicker, and operates a microgrid in a stable way compared to stand alone rule-based and reinforcement learning techniques. The findings demonstrate that deterministic safety rules with adaptive reinforcement learning is an effective and helpful approach to managing energy in smart microgrids.
S. Sreekanth, P. Kiran· International Conference on...· 0 citations
As multi-agent systems (MASs) expand in scale or dimensionality, ensuring efficient real-time distributed optimisation and consensus control becomes increasingly demanding, particularly when operating under constrained resources. To address this, the paper proposes and applies a distributed optimisation control scheme driven by bio-inspired self-triggered recurrent neural network (BSTRNN). Unlike conventional RNN that rely on periodic update schemes, the proposed BSTRNN adaptively selects update instants according to variations in internal states, thereby reducing redundant computation and improving overall efficiency. The bio-inspired architecture equips recurrent neural networks with improved stability, energy-efficient operation, and adaptive flexibility, allowing agents to converge rapidly to the global optimum while maintaining resilience against disturbances and uncertainties. Rigorous theoretical analysis identifies sufficient conditions that preclude the occurrence of Zeno behavior. Numerical simulations confirm the proposed approach, highlighting its ability to ensure consensus and achieve reliable tracking in resource-constrained MASs. Finally, the feasibility of the method is demonstrated through its deployment in a high-dimensional multi-robot arm system. Note to Practitioners—In this paper, a distributed neural dynamics optimization control method for multi-agent systems is discussed and applied to the consensus tracking task of multi-robot arms. The proposed method can enhance communication reliability during agents’ cooperative motion. This is mainly because, in multi-agent cooperative tasks, each agent operates under limited network bandwidth, and maintaining continuous communication over long periods not only causes network congestion but also results in high energy consumption. The existing approaches for collaborative motion generation in multi-agent systems are mostly developed based on the time-triggered mechanism. In this context, a challenging task arises, i.e., an on-demand updating method must be developed to ensure the seamless operation of the network. Meanwhile, the real-time performance of the cooperative task can be improved through a neural network-based solution strategy. Furthermore, the physical constraints of robot joints are considered to improve the general applicability of the proposed approach. The proposed BSTRNN scheme is applicable to distributed multi-agent systems with limited communication and computation resources, including drone swarms, networked robotic systems with bandwidth constraints, and cooperative multi-robot platforms.
Zhongwen Cao, Zhijun Zhang· IEEE Transactions on Automat...· 0 citations
Massive MIMO systems require simultaneous optimization of energy efficiency, latency, and handover performance, yet existing approaches address these objectives in isolation across disparate parameter spaces. This paper proposes a multi-state parameter self-optimization framework that jointly optimizes across five interdependent operational states—channel, mobility, system configuration, power, and latency—using deep reinforcement learning. We formulate the problem as a multi-objective Markov decision process and implement five optimization approaches: Hybrid Action Space Reinforcement Learning, Q-Learning with Kalman Filter prediction, LSTM Autoencoder for PAPR reduction, bio-inspired Integrated Fruit Fly Salp Swarm Optimization for power allocation, and a proposed Multi-Agent Deep Q-Network (MA-DQN) with experience replay. Simulation results across antenna configurations from 16 to 256 elements and user counts from 5 to 40 show that the proposed MA-DQN achieves a composite performance score of $83 \pm 1.8 / 100$ across all five states (averaged over 10 seeded runs), outperforming the best single-objective method by $\mathbf{2 6} \boldsymbol{\%}$. The framework delivers 29-73% energy efficiency improvement over fixed baselines, with the learned policy favoring moderate power (0.1-0.5W) and lower antenna counts (16-32)—consistent with analytical models that show circuit power dominance at high antenna counts.
Madhu Kumari Ray, Sasmita Mohapatra, C. J.· 2026 11th International Conf...· 0 citations