Skip to content

Deep Reinforcement Learning Method Based on Adaptive Constraint Processing

Jul 2026 · International journal of pattern recognition and artificial intelligence · 0 citations

Abstract

In the continuous decision-making process of complex engineering systems, it is often necessary to simultaneously consider long-term performance optimization and instantaneous physical executability. Therefore, the present study proposes a constraint-aware deep reinforcement learning method, CTD3, for the energy management of multi-stack fuel cell hybrid trams, aiming to coordinate long-term constraints with single-step action executability. Based on the multi-energy coupled dynamic structure, stack availability, minimum stable power, and energy storage boundaries are incorporated into the corresponding time-varying feasible domain, and the stack-level operating environment is explicitly modeled. On this basis, the energy management problem is formulated as a constrained Markov decision process (CMDP), and an adaptive Lagrangian constraint adjustment mechanism is introduced into the twin delayed deep deterministic policy gradient (TD3) framework to achieve long-term optimization and constraint maintenance requirements. Action projection and feasible power mapping execution layers are designed to convert system-level continuous actions into executable instructions that satisfy the physical boundaries at the stack level. Simulation experiments are conducted under typical urban line conditions based on a joint Python and MATLAB/Simulink platform. The results indicate that, compared with the standard TD3 algorithm, CTD3 exhibits faster convergence, improved training stability, and a decreasing trend in the overall violation rate of comprehensive constraints. Relative to the finite state machine (FSM), the equivalent hydrogen consumption of CTD3 is reduced by 6.31% and the comprehensive constraint violation rate is maintained within the 2% threshold. Meanwhile, the proposed method can guide the formation of differentiated power allocation among multiple stacks and promote the coordinated distribution of power batteries and supercapacitors. The results demonstrate that CTD3 achieves good overall performance in reducing equivalent hydrogen consumption, controlling the comprehensive constraint violation rate, and coordinating multi-source power allocation under the tested typical operating conditions, thereby providing a feasible constraint-aware reinforcement learning approach for constrained continuous control problems such as energy management of multi-stack fuel cell hybrid trams.

View source