Skip to content
Open access

Research on Optimal Power Grid Scheduling Based on Transfer Reinforcement Learning

Aug 2026 · Advanced Electromagnetics · 0 citations

Abstract

To enhance power grid adaptability amid rising renewable energy integration, this paper proposes M3-PPO, a meta-reinforcement learning algorithm that enables efficient the strategy transfer and rapid adaptation across tasks with varying energy mixes. Built upon a base framework (M-PPO) that integrates PPO and MAML, M3-PPO introduces two key innovations to overcome MAML’s training instability: a Mamba-based context encoder for richer task representation in the inner loop, and a global-local momentum update mechanism for smoother meta-parameter optimization in the outer loop. Experiments on the Grid2Op platform demonstrate that M3-PPO significantly outperforms baseline algorithms in generalization and scheduling efficiency, achieving robust performance even when simulating complex energy environments. The approach is particularly suitable for integration with antenna-enabled smart grid monitoring, wireless data acquisition, and edge-computing platforms, providing an engineering-oriented solution for adaptive, real-time, and robust power grid scheduling in modern renewable-rich energy systems.

Read PDF