Skip to content

Reinforcement learning-based adaptive control strategies for sustainable production systems

2026 · Materials Research Proceedings · 0 citations

Abstract

Abstract. The need to achieve sustainable production has become an urgent necessity in the conditions of stricter environmental requirements and the rise in the cost of energy worldwide. The classical proportionalintegralderivative controllers and linear Model Predictive Controllers are conventional model-based control strategies that by nature rely on precise process models and fixed optimization horizons and hence are not well suited to the dynamic, non-linear, and complex nature of the modern production environment. The current paper suggests a new adaptive control system based on Reinforcement Learning (RL), where the overall production system is optimized and controlled to achieve Sustainable Production Systems, with the multi-objective rewarding function, which aims to minimize the Overall Equipment Effectiveness (OEE), defect rate, specific energy consumption, and carbon dioxide emissions, and a physics-informed Digital Twin safety filter that stops unsafe policy execution in training and deployment. The proposed framework is evaluated on a multi-machine flexible manufacturing cell benchmark, which includes CNC milling, turning, and robotic assembly, and yields an OEE of 93.6, a defect rate of 0.9, a decrease in specific energy consumption of 27.8, and a decrease in CO 2 emissions of 23.4 compared to PID baseline controllers. Experiments with ablation prove all the above-mentioned in the necessity of each of the architectural constituents, and the Pareto frontier analysis proves that the proposed SAC agent is the best in the OEE-versus-energy trade-off space compared to all of the competing methods.

View source

Similar papers

2026

Reinforcement learning-based decision making for sustainable manufacturing operations

Abstract. Sustainable manufacturing involves being able to optimize productivity, energy efficiency, and environmental impact simultaneously given dynamic and uncertain operating conditions. The traditional optimization methods are unadaptable and cannot easily reflect the real-time changes in the system. This paper provides a sophisticated reinforcement learning (RL)-based decision-making model of sustainable manufacturing process. The manufacturing system is modelled as a Markov Decision Process (MDP) and a Deep Q-Network (DQN) is used to learn about the optimal control policies by interacting with the environment continuously. Multi-objective reward function is created to include production rate, energy usage, machine usage and minimization of waste. The suggested framework is tested in a virtualized smart factory setting, where the demand is stochastic and machines have variability. Comparative analysis shows that the RL-based approach outperforms the rule-based and heuristic strategies and reports remarkable energy efficiency and operational sustainability. The findings prove RL as a potential solution to adaptive and intelligent manufacturing control.

A. Jain · 0 citations
Open access Aug 2026

Optimizing production planning and control: state space design in reinforcement learning

Modern production systems are increasingly challenged by volatile and complex market conditions. To maintain their competitiveness, the application of Reinforcement Learning (RL) in production planning and control has shown considerable promise. However, the state space in RL approaches is often excessively large and complex, which triggers the curse of dimensionality and hinders learning efficiency. A significant research gap exists regarding systematic methodologies for defining, evaluating, and iteratively optimizing the state space to improve both the transparency and learning performance. This paper addresses this gap by proposing a novel methodology consisting of two steps. First, an initial state space is systematically derived from corporate planning goals with the help of a feature map. Second, Shapley Additive Explanations (SHAP) values are utilized to quantify the influence of each state feature on the RL agent’s action selection and to iteratively improve the state space accordingly. The proposed methodology is validated using a real-world industrial use case from the machinery industry. The results indicate that the methodology successfully identifies dominant state features, leading to an improved learning behavior and enhanced transparency in the RL agent’s decision-making process.

Marc Wegmann, Tobias Pfrang, Julian Stang et al. · 0 citations
Preprint Jul 2026

Scalable Supervisory HVAC Control for Linear Objectives

Advanced control of heating and cooling systems can substantially reduce energy costs and pollution. However, real-world adoption of popular algorithms among researchers, such as model predictive control (MPC) and reinforcement learning (RL), remains limited due in part to their high deployment and commissioning costs. Here, we develop two nearly commissioning-free controllers tailored to objectives that depend linearly on the controlled thermal load, such as energy costs and pollution. The controllers require at most two thermal parameters. In representative heating simulations, controller performance is robust to large parameter specification errors, suggesting potential for deployment with no tuning. The controllers maintain good occupant comfort while achieving 43 to 98% (depending on the electricity pricing and controller variant) of the performance improvement achieved by an omniscient policy with perfect model information and forecasts. These results suggest that simple, structure-exploiting controllers may capture most of the attainable value of advanced control while avoiding the data, modeling, tuning, and computational burdens that can arise with conventional MPC or RL.

W. G. Dierking, Arash Khabbazi, L. D. Premer et al. · 0 citations
2026

Machine learning–based predictive control for energy-efficient manufacturing systems

Abstract. The fact that operational costs and the environmental impact are increasing is what has made energy consumption in manufacturing systems a serious issue. In this paper, a machine learning (ML)-based predictive control model is introduced to enhance energy efficiency in the contemporary manufacturing settings. The offered solution combines predictive models based on data and Model Predictive Control (MPC) to optimize the performance of the systems in real time. Machine learning algorithms are used to predict the energy demand, process dynamics, and disturbances, as well as to make decisions proactively. Industrial case studies confirm the validity of the framework, showing great progress in terms of energy efficiency, productivity, and stability of operations.

Prabhakara Rao Kapula · 0 citations
Conference Aug 2026

A Hierarchical Hybrid Rule-Based and SARSA Reinforcement Learning Framework for Safe Energy Management in Renewable Microgrids

Microgrids must efficiently manage energy under uncertainties in renewable generation and load demand to ensure reliable and cost-effective operation. This paper investigates a microgrid system that involves renewable energy through the photovoltaic system, wind system, battery energy storage, and local load requirement with a centralized energy management system. The inflexible nature of traditional rule-based and optimizationbased approaches to solving problems can often create issues in reflection to dynamic operating conditions, and reinforcement learning approaches can produce unsafe control behavior in exploration stages. To address these issues, this paper suggests a hierarchical hybrid energy management structure that will integrate rule-based supervision and a SARSA reinforcement learning controller. Supervisory layer ensures that the system is safe by ensuring that there are operational limits such as battery state of charge limits as well as power balance conditions. The learning agent on the other hand optimizes the control choices to reduce operational costs and grid energy consumption. The outcomes of the simulation indicate that the suggested approach saves more money, learns quicker, and operates a microgrid in a stable way compared to stand alone rule-based and reinforcement learning techniques. The findings demonstrate that deterministic safety rules with adaptive reinforcement learning is an effective and helpful approach to managing energy in smart microgrids.

S. Sreekanth, P. Kiran · 0 citations