Skip to content

Author

Baomin Fang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Physics-Informed Distributionally Robust Multi-Agent Reinforcement Learning for Coordinated New-Type Power System Operation

High renewable penetration and large-scale green hydrogen production are accelerating the formation of the new-type power system (NTPS), in which electrical dispatch, electrolysis, hydrogen storage, fuel-cell reconversion, and flexible demand must be coordinated under nonlinear network physics and uncertain renewable, load, and hydrogen-demand trajectories. This study develops a physics-informed distributionally robust multi-agent reinforcement learning (PI-DRO-MARL) framework for coordinated NTPS operation with integrated electricity–hydrogen coupling. The operational objective is to minimize worst-case expected operating cost, including generation and grid-exchange cost, electrolysis and hydrogen-delivery cost, storage degradation, renewable curtailment, and load- or hydrogen-shedding penalties, while satisfying AC power-flow balance, voltage limits, line-loading limits, ramping limits, battery state-of-charge constraints, hydrogen-storage dynamics, and electrolysis/fuel-cell conversion constraints. The framework embeds physics-informed residuals and projection operators into a centralized-training decentralized-execution architecture; represents renewable, electrical-load, hydrogen-demand, and price uncertainty through statistically calibrated Wasserstein ambiguity sets; and trains agents with robust value estimation and feasibility-aware action correction. Validation is conducted on a modified IEEE 33-bus distribution network coupled with a 12-node hydrogen system, with additional scalability checks on modified IEEE 69-bus and IEEE 123-node reference systems. Across ten random seeds, the primary case shows an operating cost of USD 8850 with a 95% confidence interval of USD 8770–8940, a mean constraint-violation rate of 0.37%, and a shifted-scenario cost increase of 12.6%, outperforming deterministic optimization, stochastic programming, standard reinforcement learning (RL), proximal policy optimization (PPO), soft actor–critic (SAC), multi-agent deep deterministic policy gradient (MADDPG), constrained RL, safe RL, and robust RL baselines. Ablation, Wasserstein-radius, time-step, and stress-test analyses further show that distributional robustness, physics-informed projection, and multi-agent coordination provide distinct and complementary benefits. The results support PI-DRO-MARL as a simulation-validated architecture for real-time, uncertainty-aware NTPS dispatch, while field deployment still requires digital-twin calibration, hardware-in-the-loop testing, and site-specific operational validation.

Fei Liu, Outing Zhang, Jun Yin et al. · 0 citations