HVAC systems represent a major share of building energy consumption. Traditional control strategies are limited in coordinating energy-comfort tradeoffs across multiple zones simultaneously. Reinforcement learning (RL) offers adaptive, data-driven control that optimizes performance over time. However, deploying learned neural network controllers in safety-critical building systems remains challenging due to lack of formal safety guarantees. We propose a safety-certified deep RL framework for multi-zone residential HVAC control. Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) agents are trained in an EnergyPlus/Sinergym simulation to minimize energy consumption while maintaining thermal comfort. Post-training safety certification is performed on the PPO policy using Lipschitz-based forward invariance analysis, building on existing tools for the computation of Lipschitz constants for neural networks, to guarantee constraint satisfaction. Both agents are evaluated over an annual simulation cycle in an eight-zone variable refrigerant flow (VRF) testbed. The PPO agent achieves 67\% comfort violation reduction compared to rule-based control, while the SAC agent achieves 27.6\% energy savings. The PPO policy satisfies formal safety certification with a margin of $2.003^\circ$C. These results demonstrate the feasibility of combining reinforcement learning with post-training safety verification for multi-zone building control.
Rising electricity demand and high energy costs, particularly in rapidly growing regions, highlight the need for more efficient building energy management. Air conditioning (AC) systems account for a substantial portion of electricity consumption, yet conventional fixed-temperature control lacks adaptability to varying environmental conditions. This study introduces a Smart Air Conditioning Management System based on a Deep Q-Network (DQN) agent capable of dynamically balancing energy use and thermal comfort. The system integrates IoT-enabled environmental sensing and actuation, using ESP32 and Raspberry Pi microcontrollers, along with user feedback for automated control of AC units. Initial training and evaluation were conducted in recreated classroom-laboratory environments using simulations in EnergyPlus, and experimental deployment validated performance in real rooms. Simulation results show up to 81% energy reduction under low comfort prioritization, while field deployment achieved approximately 60% savings. These findings demonstrate that reinforcement learning enables adaptive AC control, offering a scalable approach to energy-efficient building management.
Jason Harvey Lorenzo, Justin Kyle O. Ricafort, E. Q. Macabebe· IOP Conference Series: Earth...· 0 citations
Simulation results show that the proposed Human-Interactive Lagrangian SAC (HI-LSAC) achieves significantly lower voltage-violation severity and reduced power losses compared with the other baseline methods.
An intelligent Model Predictive Control framework for optimal power flow management in microgrids, with the objective of enhancing operational resilience, reducing diesel fuel consumption, and preventing blackouts through coordinated electric vehicle (EV) charging and discharging is proposed.
H. Taha, Ahmed Abdelrahman, A. Mammeri· Energy Efficiency· 0 citations
This paper proposes an advanced metering infrastructure (AMI)-informed hierarchical energy management framework for coordinated operation of electric vehicles (EVs), photovoltaic (PV) systems, and battery energy storage systems (BESS) in campus microgrids. The proposed two-layer architecture integrates a soft actor–critic (SAC) deep reinforcement learning (DRL) agent in the upper layer with a receding horizon model predictive control (MPC) optimizer in the lower layer. The key novelty is an AMI-to-control pipeline that transforms historical 15 min smart-meter measurements into operational flexibility features and embeds them into a hierarchical SAC–MPC architecture, where the DRL layer provides adaptive coordination and the MPC layer enforces grid, storage, and EV-service constraints. The proposed framework using the real-world Pecan Street data (15 min resolution) of 73 homes across Austin, Texas and California (2014–2019) achieves a 53.1% cost reduction and a 25.7% peak demand reduction when compared with uncontrolled charging, and the proposed framework outperforms MPC-only (50.9%), DRL-only (−5.2%), and rule-based (5.1%) baselines. The statistically significant contributions of network-aware constraints, demand-response activation, and predictive look-ahead horizon are statistically significant (n = 10 independent runs) contributions (p = 0.001). The state representation informed by AMI offers directional cost improvement (+8.4%, p = 0.055) with 11% faster convergence of training. The zero network constraint violation is observed in all evaluation scenarios and the average MPC solve time is around 150 ms, which is much less than the 15 min sampling period. Sensitivity analyses show that the hierarchical DRL–MPC architecture remains computationally feasible across EV penetration, seasonal, and forecast-uncertainty scenarios. However, BESS provided no net economic benefit under the evaluated energy-only TOU tariff, increasing weekly cost by $15.25 and peak grid demand by 14.2 kW. Break-even analysis indicates that demand charges of approximately $9.9/kW per month are required for BESS to become cost-effective in the proxy system, highlighting that storage value depends strongly on tariff design and peak-demand objective formulation.
Mousa A. Aljabri, Mohammed O. Bahabri, Nasser A. Alakhrash et al.· Energies· 0 citations
Rising global energy demand and increasing decarbonization requirements have intensified the need for intelligent building energy management capable of handling nonlinear dynamics and multi-objective operational trade-offs. Conventional discrete-time and simulation-dependent control strategies often struggle to maintain temporal continuity, adaptive responsiveness, and consistent performance across heterogeneous building environments. Addressing these limitations, NODE-RL-BEM (Neural Ordinary Differential Equation Reinforcement Learning for Building Energy Management) introduces a unified continuous-time optimization paradigm that jointly models system dynamics and learns adaptive control policies. The approach integrates heterogeneous operational data, temporal state embeddings, neural differential equation modeling, and multi-objective reinforcement learning within a cohesive architecture designed for predictive and responsive energy optimization. Performance evaluation conducted on the ASHRAE Great Energy Predictor III dataset and the Intelligent Indoor Environment Dataset demonstrates the effectiveness of the proposed framework, achieving 42-48% energy savings, maintaining comfort violations below 0.5%, and improving indoor air quality by 28-35%. The framework further achieves a generalization score of 0.91 across diverse building operational scenarios, confirming strong transferability and stability. Continuous-time dynamics learning improves predictive fidelity and ensures smooth state evolution, while adaptive reinforcement learning enables robust decision-making under dynamic environmental and occupancy variations. Scalable applicability to multi-zone building environments highlights practical deployment feasibility. This work establishes a novel continuous-time dynamic-policy learning paradigm that integrates predictive modeling with real-time adaptive control, advancing data-driven intelligent building operation toward sustainable and autonomous energy management.
Hamoud H. Alshammari· Scientific Reports· 0 citations
The rapid growth of energy markets and demand-side response programs has created a significant need for intelligent building-level control strategies to capture high volatility energy consumption in response to price signals and grid conditions. This paper presents a reinforcement learning (RL)-based building energy management framework that models commercial buildings as active, grid-interactive assets capable of providing real-time flexibility while maintaining occupant comfort. The proposed approach integrates historical and real-time data from IoT sensors, HVAC systems, and weather forecasts to build an adaptive environment for RL agents. The RL model is applied to learn optimal control policies that minimize operational energy cost in response to flexibility markets through load shifting, peak shaving, and short-term demand response actions. The framework also incorporates a forecasting module to predict 30-minute interval energy consumption using deep learning, enabling proactive decision-making under uncertainty. Results from simulation experiments demonstrate that the RL agents achieve significant cost savings compared to rule-based control strategies and offer a reliable, automated control to unlock underlying flexibility within building systems. The paper discusses problem formulation, algorithmic development, simulation workflows, comparative metrics, and practical deployment considerations for Saudi Arabia's smart city initiatives.
Abdulaziz Almalaq· 2026 6th International Confe...· 0 citations