INTELLIGENT SCHEDULING OF PV–STORAGE–CHARGING INTEGRATED STATIONS VIA GTRXL-PPO WITH CURRICULUM LEARNING
To address the long-horizon sequential decision-making task, characterized by complex temporal dependencies, non-stationary dynamics, and high stochasticity in distribution-level PV–storage–charging systems, this paper develops a deep reinforcement learning framework that combines Gated Transformer‑XL (GTrXL) with Proximal Policy Optimization (PPO). Cross-segment memory captures long‑range temporal dependencies and time‑aware encodings reinforce intraday periodicity. Training adopts a five‑stage curriculum with adaptive KL control and auxiliary multi‑task heads to improve sample efficiency. A group‑normalized, potential‑based reward unifies economic performance, grid friendliness, and storage health. In simulation, the agent learns a structured six-phase daily policy and achieves a 94.4% charging completion rate and 92.8% PV utilization, reduces average daily electricity purchase cost by 15%, and keeps grid peak power within a 60 kW soft limit. Across five seeds, returns improve by 78.4% over a feed-forward PPO baseline and by 23.6% over a vanilla GTrXL-PPO, demonstrating the benefits of long-memory RL for coordinated PV–storage–charging operation. The framework enforces feasibility via continuous action mapping with ramp-rate/jerk constraints and supports millisecond-level inference. Uncertainty-aware shaping improves robustness; gains are statistically significant across five seeds via paired tests.