Skip to content

Safe Multi-Agent Collaborative Learning for Networked Grid Operation Under Power Network Coupling Constraints

Jul 2026 · International journal of pattern recognition and artificial intelligence · 0 citations

TL;DR

This paper formulate networked grid operation as a constrained decentralized partially observable Markov decision process and proposes a safe multi-agent collaborative learning framework that aims to reduce operating cost, load shedding, renewable curtailment, and carbon-relevant corrective burden.

Abstract

Modern grid operation is increasingly a sequential collaborative control problem under renewable uncertainty, storage dynamics, flexible demand, transmission coupling, and carbon-aware corrective redispatch. This paper focuses on sub-hourly dynamic OPF assistance and safe regional redispatch after forecast updates or emergent network stress. We formulate networked grid operation as a constrained decentralized partially observable Markov decision process and propose a safe multi-agent collaborative learning framework. The method integrates dynamic transfer-stress tracing, a dual-view graph encoder over the physical grid and a real-time stress graph, consensus-based dual coordination for globally coupled constraints, and a differentiable safety projection that maps tentative decisions to executable actions. The framework aims to reduce operating cost, load shedding, renewable curtailment, and carbon-relevant corrective burden, while maintaining feasibility with respect to local linearized surrogate limits and empirically reducing violations in the full simulator. Experiments on chronological grid benchmarks evaluate comparison, ablation, robustness, efficiency, statistical evidence, visualization, and generalization, showing improved trade-offs among cost, reliability, safety, and renewable accommodation.

View source

Similar papers

2026

Collaborative Task Offloading in Space Computing Power Network: A World Model-Based Multi-Agent Reinforcement Learning Approach

Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.

Yuqi Cong, Zhiwei Wei, Jiarui Chen et al. · 0 citations
Open access Jul 2026

Physics-Informed Distributionally Robust Multi-Agent Reinforcement Learning for Coordinated New-Type Power System Operation

High renewable penetration and large-scale green hydrogen production are accelerating the formation of the new-type power system (NTPS), in which electrical dispatch, electrolysis, hydrogen storage, fuel-cell reconversion, and flexible demand must be coordinated under nonlinear network physics and uncertain renewable, load, and hydrogen-demand trajectories. This study develops a physics-informed distributionally robust multi-agent reinforcement learning (PI-DRO-MARL) framework for coordinated NTPS operation with integrated electricity–hydrogen coupling. The operational objective is to minimize worst-case expected operating cost, including generation and grid-exchange cost, electrolysis and hydrogen-delivery cost, storage degradation, renewable curtailment, and load- or hydrogen-shedding penalties, while satisfying AC power-flow balance, voltage limits, line-loading limits, ramping limits, battery state-of-charge constraints, hydrogen-storage dynamics, and electrolysis/fuel-cell conversion constraints. The framework embeds physics-informed residuals and projection operators into a centralized-training decentralized-execution architecture; represents renewable, electrical-load, hydrogen-demand, and price uncertainty through statistically calibrated Wasserstein ambiguity sets; and trains agents with robust value estimation and feasibility-aware action correction. Validation is conducted on a modified IEEE 33-bus distribution network coupled with a 12-node hydrogen system, with additional scalability checks on modified IEEE 69-bus and IEEE 123-node reference systems. Across ten random seeds, the primary case shows an operating cost of USD 8850 with a 95% confidence interval of USD 8770–8940, a mean constraint-violation rate of 0.37%, and a shifted-scenario cost increase of 12.6%, outperforming deterministic optimization, stochastic programming, standard reinforcement learning (RL), proximal policy optimization (PPO), soft actor–critic (SAC), multi-agent deep deterministic policy gradient (MADDPG), constrained RL, safe RL, and robust RL baselines. Ablation, Wasserstein-radius, time-step, and stress-test analyses further show that distributional robustness, physics-informed projection, and multi-agent coordination provide distinct and complementary benefits. The results support PI-DRO-MARL as a simulation-validated architecture for real-time, uncertainty-aware NTPS dispatch, while field deployment still requires digital-twin calibration, hardware-in-the-loop testing, and site-specific operational validation.

Fei Liu, Outing Zhang, Jun Yin et al. · 0 citations
Open access Jul 2026

Data-Driven Risk-Aware Approximate Dynamic Programming Algorithm for Resilient Power System Operation Under High Renewable Uncertainty

The accelerating integration of renewable energy sources into modern power grids has created unprecedented operational challenges, with significant system cost volatility under extreme uncertainty events. To address this challenge, this paper presents a risk-aware stochastic approximate dynamic programming (SADP) algorithm based on machine learning and parallel computing architectures. The algorithm learns optimal coordination strategies for source-grid-load-storage resources while explicitly quantifying and mitigating tail risk events that conventional approaches overlook. First, a risk-averse stochastic optimization model is constructed, which captures the complex interdependencies between renewable generation uncertainty, demand variability, and flexible resource coordination through second-order cone programming formulations. This model integrates the GlueVaR (Glued Value-at-Risk) metric, enabling simultaneous optimization across multiple risk horizons with adjustable conservatism parameters. Second, to solve the established model efficiently, an SADP algorithm based on risk-averse approximate value functions (RAVFs) is proposed, in which the training process of the RAVFs employs machine learning principles to directly encode risk preferences into operational decisions. By integrating GlueVaR into offline training across 5000 probabilistically weighted scenarios, the algorithm discovers emergent coordination patterns between distributed resources, which are rarely identified by human operators. Third, a large-scale parallel computing architecture is implemented for the SADP algorithm. This architecture decomposes the multi-period optimization problem into single-period coordinated sub-problems. During offline training, parallel computing of a series of single-period sub-problems can be performed across all probabilistic scenarios, significantly reducing training time. Extensive validation on both the modified IEEE 33-bus and 69-bus systems with integrated wind turbines, photovoltaic plants, energy storage systems, and demand response capabilities demonstrates remarkable performance improvements. Convergence analysis reveals that the AVFs stabilize within 30 training iterations, achieving sub-160 s solution times in online application even for complex networks with heterogeneous resources. By enabling real-time risk-aware decision-making under severe uncertainty, the proposed method provides grid operators with actionable strategies that balance economic efficiency and operational resilience.

Zike Guo, Peng Yang, Xue Du et al. · 0 citations
Open access Jul 2026

Learning operational compatibility in distribution networks under multi-step renewable dynamics.

High penetration of renewable generation introduces rapid variability and multi-timescale uncertainty into distribution network operations, making real-time feasibility assessment increasingly difficult for traditional optimization-based dispatch frameworks. This paper proposes a learning-based operational compatibility framework that identifies whether candidate dispatch states remain feasible under Multi-Step renewable dynamics. The proposed method constructs a compatibility learning model that maps system states, renewable forecasts across multiple time horizons, and network operating conditions to a probabilistic compatibility score representing the likelihood that voltage, line flow, and power balance constraints are satisfied. A Multi-Step renewable dynamics representation is introduced to capture correlated variability across short-term and intra-day forecasting intervals, enabling the model to anticipate feasibility degradation caused by forecast uncertainty propagation. The learning architecture integrates network structural information and operational features to approximate the feasible operating manifold without repeatedly solving computationally intensive AC power flow problems. Extensive simulations on modified IEEE distribution feeders with high renewable penetration demonstrate that the proposed framework can correctly identify feasible operating conditions in approximately 94.6% of scenarios while reducing feasibility evaluation time by over 82% compared with conventional AC-OPF screening procedures. Under severe renewable fluctuation conditions, the method maintains a compatibility prediction accuracy exceeding 90%, allowing operators to rapidly filter infeasible dispatch candidates and maintain secure network operation across multiple forecast horizons.

Lujie Qi, Zhao Li, Kexin Wang et al. · 0 citations
Conference Jul 2026

Multi-Agent Reinforcement Learning Based Latency-Aware Hierarchical Control Enabling Electric-Vehicle Based Virtual Power Plants to Participate in Balancing Markets

Balancing markets require flexible resources that can promptly follow dispatch signals. Aggregated fleets of electric vehicles (EVs) operated as electric-vehicle virtual power plants (EV VPPs) are promising candidates. Aggregators must control the total power of EV chargers to track dispatch signals while satisfying individual EV users' charging demands. Conventional centralized optimization methods can achieve high tracking performance. However, they rely on global information and require solving large-scale optimization problems, which impose high computational and communication burdens and limit scalability. To address this issue, this paper proposes a two-level hierarchical control scheme based on the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. At the upper level, each charging station is modeled as an agent, and at the lower level its policy allocates charging and discharging power to the individual EV chargers. At runtime, each station-level agent uses only local observations and the broadcast dispatch signal. We present a case study on participation in Japan's balancing market Secondary 2 (S2) product. The study evaluates the controller on an EV VPP consisting of five stations with a total of 50 Level 2 chargers. The proposed method achieves a dispatch tracking rate (fraction of dispatch intervals with aggregate power inside the market-defined tracking error band) of 97 percent within the allowable tracking error band around the dispatch signal. It also achieves an 80 percent Target SoC satisfaction rate, where the Target SoC is the user-specified departure-time state of charge (SoC). Overall, this method reduces online computation time and communication latency while maintaining high tracking performance and userdemand satisfaction. These results suggest that MADDPG-based hierarchical control provides a practical control scheme for large EV fleets when latency constraints hinder centralized control.

Koshin Hayashi, Sihui Xue, Daisuke Kodaira · 0 citations
Open access Aug 2026

Edge-native intelligent scheduling for virtual power plants: A multi-scale perception and constrained reinforcement learning approach

The proliferation of distributed energy resources at the edge of distribution networks provides substantial flexibility for virtual power plant (VPP) operation. However, existing methods often rely on aggregate load information and homogeneous scheduling policies. They, therefore, overlook device-specific response characteristics, heterogeneous response times, and operational safety constraints. This paper presents EDGE-VPP, an end-to-end scheduling framework that connects fine-grained load perception with safety-aware decision-making across multiple temporal scales. At the perception layer, a Load Decomposition Transformer (LDT) uses learnable multi-frequency positional encodings and device-specific attention heads. It jointly detects appliance states and disaggregates device power from aggregate measurements. At the coordination layer, a three-tier cloud–edge–device architecture assigns sub-second emergency response to devices, minute-level economic dispatch to edge controllers, and hour-ahead planning to the cloud. Bidirectional information exchange mitigates conflicts among these control layers. At the optimization layer, multi-constraint proximal policy optimization factorizes continuous and discrete actions. Adaptive Lagrange multipliers enforce voltage and current limits, while two value estimators stabilize policy learning. Experiments on REDD, UK-DALE, and a self-constructed VPP dataset show that LDT reduces mean absolute error by up to 6.86% and improves the F1-score by 3.51% over the Transformer baseline. The complete EDGE-VPP framework also achieves the lowest operating cost and the fewest constraint violations among the evaluated scheduling methods.

Yuandong Jiang, Mingyu Ou, Jiangnan Li · 0 citations