A forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data, which preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates.
Abstract
The growing variability of renewable generation increases the need for fast and flexible grid-balancing mechanisms. Existing frameworks for distributed energy resource aggregations (DERAs) rely on short-term forecasts of net demand, making their performance highly sensitive to prediction errors. In this paper we present a forecast-free reinforcement learning (RL) framework for DERA allocation that learns optimal policies directly from operational data. We model the DERA dynamics as a deterministic linear system and the exogenous net load as a feature-based linear Markov process, capturing short-range temporal dependencies without explicit forecasting. We derive a closed-form expression for the optimal policy, which is learned through a least-squares value iteration (LSVI) algorithm using data collected across episodes. The proposed framework preserves the interpretability and constraint satisfaction of DER model while adapting to stochastic demand variations through data-driven updates. Numerical experiments on real California Independent System Operator (CAISO) net-demand data demonstrate that the learned controller achieves high tracking accuracy and stable regulation across heterogeneous DER aggregators without requiring any demand prediction.
The rapid growth of energy markets and demand-side response programs has created a significant need for intelligent building-level control strategies to capture high volatility energy consumption in response to price signals and grid conditions. This paper presents a reinforcement learning (RL)-based building energy management framework that models commercial buildings as active, grid-interactive assets capable of providing real-time flexibility while maintaining occupant comfort. The proposed approach integrates historical and real-time data from IoT sensors, HVAC systems, and weather forecasts to build an adaptive environment for RL agents. The RL model is applied to learn optimal control policies that minimize operational energy cost in response to flexibility markets through load shifting, peak shaving, and short-term demand response actions. The framework also incorporates a forecasting module to predict 30-minute interval energy consumption using deep learning, enabling proactive decision-making under uncertainty. Results from simulation experiments demonstrate that the RL agents achieve significant cost savings compared to rule-based control strategies and offer a reliable, automated control to unlock underlying flexibility within building systems. The paper discusses problem formulation, algorithmic development, simulation workflows, comparative metrics, and practical deployment considerations for Saudi Arabia's smart city initiatives.
Abdulaziz Almalaq· 2026 6th International Confe...· 0 citations
Dynamic electricity tariffs are increasingly deployed to manage demand-side flexibility in decarbonising power systems, making AI-driven load forecasting a critical enabler of adaptive pricing. However, existing studies largely treat forecasting and tariff design as independent problems, evaluating models on predictive accuracy alone while neglecting the feedback effects through which price signals reshape consumer behaviour and introduce non-stationarity into the very demand distributions forecasts depend upon. This gap leaves practitioners without coherent guidance on how to structure, govern, and adapt forecasting models in price-responsive environments. This paper addresses the gap through a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA)-informed structured review of peer-reviewed studies spanning statistical, machine learning, deep learning, probabilistic, and reinforcement learning forecasting paradigms. Three principal findings emerge. First, dynamic pricing fundamentally invalidates static forecasting assumptions by inducing a distributional shift in demand data, making robustness to behavioural feedback a first-order requirement for deployment. Second, no single forecasting paradigm simultaneously satisfies the requirements of accuracy, interpretability, uncertainty quantification, and regulatory defensibility that dynamic tariff systems impose; layered, role-specific architectures are therefore operationally necessary. Third, explainability and governance constraints are structural requirements for tariff-oriented forecasting, not optional enhancements, because forecast outputs directly influence economically and socially consequential pricing decisions. Building on these findings, the paper proposes a Forecast-Driven Dynamic Tariff Design Framework that separates forecasting intelligence from pricing authority, embeds uncertainty management as a first-class design element, and positions a governance layer as the mandatory interface between predictive outputs and consumer-facing tariff signals. The framework provides a practical and regulatorily defensible foundation for deploying adaptive, resilient, and equitable electricity tariffs in data-intensive power systems.
Oluwagbenga Apata, Mukovhe Ratshitanga, I. Davidson· Energies· 0 citations
Driven by artificial intelligence and cloud computing, hyperscale data centers are becoming one of the fastest-growing electrical loads worldwide and are increasingly recognized as a new class of flexible loads capable of supporting demand response (DR) and the integration of variable renewable energy (VRE). However, their two principal control levers—IT workload scheduling and cooling system operation—have traditionally been managed in a decoupled manner, leaving both energy efficiency and demand-side flexibility under-exploited. This paper proposes a deep reinforcement learning (DRL) framework that jointly co-schedules computing and thermal resources so that a hyperscale data center can operate as a grid-interactive flexible load. We formulate the joint problem as a constrained Markov Decision Process and develop an actor-critic algorithm combining Deep Deterministic Policy Gradient with a safety shield mechanism to guarantee thermal constraint satisfaction during both training and deployment. A high-fidelity digital twin simulation environment enables safe Sim-to-Real training. Extensive experiments demonstrate that the proposed approach reduces total electricity consumption by 18-25% compared to baseline controllers, cuts thermal violations by over 90%, and maintains service level agreement compliance, while broadening the controllable power envelope of the facility to provide a technical basis for participating in DR programs and aligning data-center power profiles with renewable generation. The framework bridges IT-side and facility-side control and supports the evolution of hyperscale data centers from passive electricity consumers toward active, grid-interactive participants in renewable-penetrated power systems.
It is demonstrated that lightweight aggregation strategies can substantially improve empirical safety in federated reinforcement learning while preserving standard communication protocols.
This paper presents a bilevel Reinforcement Learning (RL) framework for optimizing Electric Vehicle (EV) charging through price-mediated coordination between grid operators and charging stations. Unlike prior work relying on direct control or manual subgoal engineering, the proposed approach uses dynamic pricing as an implicit coordination signal to address a complex multi-objective optimization problem involving grid stability, user satisfaction, and economic efficiency. To manage this complexity, the problem is decomposed into two levels comprised of an upper-level Distribution System Operator (DSO) that determines dynamic pricing strategies, and multiple lower-level Load Aggregators (LAs) responsible for EV charging decisions at individual stations in response to these prices. This bilevel structure captures the leader–follower interaction between DSOs and LAs, with each level operating at different temporal scales. Deep Deterministic Policy Gradient (DDPG) agents are deployed at both levels, enabling adaptive decision-making under operational constraints. Extensive simulations compare the framework against multiple Rule-Based Control (RBC) baselines. Results demonstrate that the DDPG-based DSO achieves a 42.4% higher mean reward and 19.1% higher profit compared to the best-performing RBC baseline, while preserving grid stability and user satisfaction. These results validate the effectiveness of bilevel RL for complex energy optimization problems, highlighting its potential as a scalable control paradigm for smart management systems.
D. Vamvakas, Christos D. Korkas, E. Kosmatopoulos· Energies· 0 citations
This paper proposes H-UPF (Hybrid Universal Policy with Forecasting), a hybrid intelligent framework for scalable sequential decision-making in heterogeneous environments under uncertainty. The architecture integrates probabilistic multi-horizon forecasting via a Temporal Fusion Transformer with continuous control via Proximal Policy Optimization, embedding predictive quantile distributions directly into the agent’s state representation. A Dynamic Adaptation Layer normalizes observations relative to instance-specific scales, enabling zero-shot policy transfer across environments with 18.5× variability in operating characteristics — without inter-agent communication or per-instance retraining. Validated on two real-world residential energy management datasets (REFIT: 20 UK households; CityLearn: 6 US buildings with real PV profiles), the framework achieves 88.4% of the theoretical optimum in zero-shot transfer, outperforming meta-learning (MAML-PPO) by 8.4 percentage points (Wilcoxon p = 0.003, Cohen’s d = 1.42). Ablation analysis identifies the adaptation layer as the dominant contributor (−16.2 p.p. upon removal), while probabilistic forecasting adds +6.8 p.p. through proactive scheduling. The learned policy is robust to reward parameter variations (≤3.2 p.p. sensitivity across 5× range) and supports practical deployment: 9.8 h one-time training, 18.4 ms inference per control step.
A. Tokhmetov, L. Tanchenko, M. Kenesbai· Bulletin of Manash Kozybayev...· 0 citations