Skip to content
Open access

Fusion of adaptive decoupling control and deep reinforcement learning for intelligent waste heat recovery

Aug 2026 · Journal of Vibration and Control · 0 citations · 33 references

TL;DR

The proposed NMMC-DRL strategy is evaluated through comprehensive simulations and physical experiments on a Roots-expander-based waste heat recovery platform and shows effectiveness in improving disturbance rejection, tracking accuracy, and operational robustness.

Abstract

Efficient recovery of low-grade industrial waste heat is important for improving energy utilization; however, Roots-expander-based waste heat recovery systems remain difficult to control because of strong multi-parameter coupling, nonlinear operating characteristics, model-switching transients, and stochastic gas-source disturbances. To address these challenges, this study proposes an intelligent hybrid control strategy, termed NMMC-DRL, for a Roots-expander-based waste heat recovery system. The main novelty of the proposed method is that deep reinforcement learning (DRL) is not used as a standalone controller, but as an upper-level residual compensation mechanism embedded in a nonlinear multi-model adaptive decoupling control (NMM-ADC) framework. In this architecture, the lower-level NMM-ADC controller handles nominal nonlinear coupling dynamics and preserves the stabilizing role of the model-based control structure, whereas the upper-level DRL agent, trained using the deep deterministic policy gradient algorithm, learns real-time compensatory actions for unmodeled dynamics, coupling residuals, model-switching effects, and unknown gas-source disturbances. Furthermore, an extended state representation and a compensation-oriented reward function are designed to incorporate tracking errors, coupling residual information, and switching-related effects into the learning process, enabling smoother and more adaptive compensation. The proposed NMMC-DRL strategy is evaluated through comprehensive simulations and physical experiments on a Roots-expander-based waste heat recovery platform. Under random gas-source fluctuations, the controller limits speed variations to within ±2.5%, reduces settling time by approximately 40%, and decreases overshoot by approximately 60%. These results demonstrate the effectiveness of the proposed hybrid control strategy in improving disturbance rejection, tracking accuracy, and operational robustness.

Read PDF

Similar papers

Open access Jul 2026

AI-Enhanced Chemical Separation Based on Deep Learning for Intelligent Process Control

Chemical process operations require fault-adaptive and energy-aware automation, yet conventional control methods often struggle to maintain robustness, fault tolerance, and energy efficiency under dynamic operating conditions. This paper presents an artificial intelligence-enhanced process monitoring and control framework for the Tennessee Eastman process that integrates a hybrid temporal convolutional neural network-bidirectional gated recurrent unit encoder for fault detection with a safety-constrained reinforcement learning-based adaptive controller. The proposed framework achieves superior fault detection performance, with an F1 score of 0.957, an area under the receiver operating characteristic curve of 0.981, and a median detection time of 10.2 s, outperforming long short-term memory and autoencoder baseline models. It also delivers improved control performance, achieving an energy index of 0.866 and a control efficiency of 1.113, while reducing energy consumption compared with the long short-term memory, autoencoder, and model predictive control baselines. These findings demonstrate that combining temporal feature learning with safety-constrained policy optimization provides a practical approach for developing more resilient, energy-efficient industrial process automation systems.

Nannan Li · 0 citations
Preprint Aug 2026

Safe Deep Reinforcement Learning for Energy-Efficient HVAC Control in Multi-Zone Residential Buildings

HVAC systems represent a major share of building energy consumption. Traditional control strategies are limited in coordinating energy-comfort tradeoffs across multiple zones simultaneously. Reinforcement learning (RL) offers adaptive, data-driven control that optimizes performance over time. However, deploying learned neural network controllers in safety-critical building systems remains challenging due to lack of formal safety guarantees. We propose a safety-certified deep RL framework for multi-zone residential HVAC control. Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) agents are trained in an EnergyPlus/Sinergym simulation to minimize energy consumption while maintaining thermal comfort. Post-training safety certification is performed on the PPO policy using Lipschitz-based forward invariance analysis, building on existing tools for the computation of Lipschitz constants for neural networks, to guarantee constraint satisfaction. Both agents are evaluated over an annual simulation cycle in an eight-zone variable refrigerant flow (VRF) testbed. The PPO agent achieves 67\% comfort violation reduction compared to rule-based control, while the SAC agent achieves 27.6\% energy savings. The PPO policy satisfies formal safety certification with a margin of $2.003^\circ$C. These results demonstrate the feasibility of combining reinforcement learning with post-training safety verification for multi-zone building control.

Oussama Ziadi, A. Rochd, S. I. Kaitouni et al. · 0 citations
Conference Jul 2026

Hybrid Deep Learning and First-Principles Model for Predictive Control of Solar Direct Steam Generation

Linear Fresnel Reflector (LFR) systems for Direct Steam Generation (DSG) exhibit strong nonlinear and transient dynamics due to diurnal solar variability, making real-time control challenging. To address this, a Gated Recurrent Unit (GRU)-based deep neural network surrogate replaces the computationally intensive first-principles LFR model and is embedded within a Nonlinear Model Predictive Control (NMPC) framework for fast prediction. Trained using scheduled sampling, the surrogate achieves prediction fits of 87.39% for steam quality and 90.55% for LFR outlet pressure. The GRU-based LFR model is integrated with a first-principles Steam Drum (SD) model in a centralized NMPC scheme to regulate steam quality and drum water level. Closed-loop simulations under varying solar insolation with cloud cover demonstrate effective tracking of time-varying setpoints, validating the proposed hybrid predictive control approach.

Dibyajyoti Baidya, Ashutosh K. Singh, M. Bhushan et al. · 0 citations
2026

Reinforcement learning-based adaptive control strategies for sustainable production systems

Abstract. The need to achieve sustainable production has become an urgent necessity in the conditions of stricter environmental requirements and the rise in the cost of energy worldwide. The classical proportionalintegralderivative controllers and linear Model Predictive Controllers are conventional model-based control strategies that by nature rely on precise process models and fixed optimization horizons and hence are not well suited to the dynamic, non-linear, and complex nature of the modern production environment. The current paper suggests a new adaptive control system based on Reinforcement Learning (RL), where the overall production system is optimized and controlled to achieve Sustainable Production Systems, with the multi-objective rewarding function, which aims to minimize the Overall Equipment Effectiveness (OEE), defect rate, specific energy consumption, and carbon dioxide emissions, and a physics-informed Digital Twin safety filter that stops unsafe policy execution in training and deployment. The proposed framework is evaluated on a multi-machine flexible manufacturing cell benchmark, which includes CNC milling, turning, and robotic assembly, and yields an OEE of 93.6, a defect rate of 0.9, a decrease in specific energy consumption of 27.8, and a decrease in CO 2 emissions of 23.4 compared to PID baseline controllers. Experiments with ablation prove all the above-mentioned in the necessity of each of the architectural constituents, and the Pareto frontier analysis proves that the proposed SAC agent is the best in the OEE-versus-energy trade-off space compared to all of the competing methods.

Apoorva Verma · 0 citations