Reinforcement learning approach for integrated dynamic pricing and inventory control
Abstract
This study focuses on the pricing and inventory coordination decision-making challenges faced by enterprises in dynamic markets. In response to the limitations of traditional methods in dealing with the interweaving of multiple time scales and the trade-offs of multiple objectives, a novel joint optimization algorithm integrating hierarchical reinforcement learning and the maximum entropy framework is proposed. The core of this method lies in constructing a dual-layer decision making architecture with close coupling between strategy and tactics. The high-level strategy acts as a meta-controller, analyzing the macro operational situation at a low frequency and outputting continuous and abstract strategic instructions to set the tone for the subsequent period's operations. The bottom layer is a flexible executor based on the principle of maximum entropy, which generates specific pricing and replenishment actions based on the detailed environmental state and the high-level instructions. The key innovation of the algorithm lies in introducing a learnable strategic alignment reward function, which is integrated into the bottom-level optimization objective, aiming to dynamically evaluate and encourage specific actions to be consistent with the high-level strategic intentions, thereby achieving an effective balance between immediate business returns, active exploration, and strategic following. Through an end-to-end joint training mechanism, the high-level and low-level strategies can evolve collaboratively. To fully validate the algorithm’s performance, systematic simulation experiments were designed. Benchmark performance comparison experiments show that the proposed algorithm significantly outperforms mainstream methods in terms of comprehensive benefits, operational efficiency, and strategic stability. Abandonment experiments confirm that the hierarchical architecture is the key to improving decision-making efficiency and convergence speed. Special mechanism tests have verified that strategic alignment rewards are indispensable for ensuring the accurate interpretation and execution of high-level instructions. Moreover, in zero-sample generalization tests in complex scenarios, the algorithm demonstrates excellent adaptability and robustness, indicating that its hierarchical abstraction can capture the general decision-making patterns across scenarios.