Physics-Informed Distributionally Robust Multi-Agent Reinforcement Learning for Coordinated New-Type Power System Operation
Abstract
High renewable penetration and large-scale green hydrogen production are accelerating the formation of the new-type power system (NTPS), in which electrical dispatch, electrolysis, hydrogen storage, fuel-cell reconversion, and flexible demand must be coordinated under nonlinear network physics and uncertain renewable, load, and hydrogen-demand trajectories. This study develops a physics-informed distributionally robust multi-agent reinforcement learning (PI-DRO-MARL) framework for coordinated NTPS operation with integrated electricity–hydrogen coupling. The operational objective is to minimize worst-case expected operating cost, including generation and grid-exchange cost, electrolysis and hydrogen-delivery cost, storage degradation, renewable curtailment, and load- or hydrogen-shedding penalties, while satisfying AC power-flow balance, voltage limits, line-loading limits, ramping limits, battery state-of-charge constraints, hydrogen-storage dynamics, and electrolysis/fuel-cell conversion constraints. The framework embeds physics-informed residuals and projection operators into a centralized-training decentralized-execution architecture; represents renewable, electrical-load, hydrogen-demand, and price uncertainty through statistically calibrated Wasserstein ambiguity sets; and trains agents with robust value estimation and feasibility-aware action correction. Validation is conducted on a modified IEEE 33-bus distribution network coupled with a 12-node hydrogen system, with additional scalability checks on modified IEEE 69-bus and IEEE 123-node reference systems. Across ten random seeds, the primary case shows an operating cost of USD 8850 with a 95% confidence interval of USD 8770–8940, a mean constraint-violation rate of 0.37%, and a shifted-scenario cost increase of 12.6%, outperforming deterministic optimization, stochastic programming, standard reinforcement learning (RL), proximal policy optimization (PPO), soft actor–critic (SAC), multi-agent deep deterministic policy gradient (MADDPG), constrained RL, safe RL, and robust RL baselines. Ablation, Wasserstein-radius, time-step, and stress-test analyses further show that distributional robustness, physics-informed projection, and multi-agent coordination provide distinct and complementary benefits. The results support PI-DRO-MARL as a simulation-validated architecture for real-time, uncertainty-aware NTPS dispatch, while field deployment still requires digital-twin calibration, hardware-in-the-loop testing, and site-specific operational validation.