Safety-Constrained Deep Reinforcement Learning for Source–Load–Storage Coordinated Operation of Green Low-Carbon Data Centers
Abstract
Green low-carbon data centers operate as coupled cyber-energy systems whose dispatch must coordinate renewable generation, grid exchange, battery storage, cooling load, flexible computing workload, carbon-intensity signals, and reliability constraints. This study develops and evaluates a safety-constrained deep reinforcement learning framework for source–load–storage coordinated operation of a grid-connected green data center. The operating problem is formulated as a constrained Markov decision process with state variables describing the IT load, deferrable workload backlog, renewable availability, electricity price, marginal carbon intensity, battery state of charge, server-room temperature, reserve margin, and calendar context. The action space covers grid import and export, renewable utilization, storage charge and discharge, workload shifting, and cooling control. The learning architecture combines a constrained actor–critic policy, adaptive Lagrangian safety critics, and a control barrier function (CBF)-based action shield that projects unsafe actions onto an explicitly defined operating set before plant execution. The shield is specified as a low-dimensional quadratic projection over state-dependent SOC, thermal, reserve, SLA, and grid-interface constraints, while cumulative risks are priced through Lagrangian safety budgets during policy training. The evaluation uses a controlled and auditable benchmark simulation with normalized public-data-compatible profiles, declared scenarios, random seeds, neural-network settings, and mechanism-matched baselines; it is not a telemetry-based verification or hardware certification of a deployed data center. Within this declared benchmark, the proposed safe DRL controller produces a simulated 13.1% emission reduction relative to the Rule-based controller, 95.8% renewable utilization, a normalized annual cost of 0.91, and fewer boundary contacts than the tested unconstrained, Lagrangian-only, and shield-only PPO variants. These percentages are simulator outputs relative to the stated benchmark and must not be interpreted as measured field savings. The results show how separating reward learning, cumulative safety pricing, and one-step engineering projection changes low-carbon dispatch within the specified model.