Skip to content
Conference

Inventory Optimisation with Deep Reinforcement Learning in Supply Chain Management

Aug 2026 · International Conference on Automation and Computing · pp. 1-6 · 0 citations · 22 references

Abstract

Inventory optimisation is critical for balancing operational costs and customer service quality. Traditional inventory methods, including EOQ, stochastic policies and MRP, depend on static demand assumptions and fixed distributions, limiting their adaptability to market volatility, seasonality and shifting customer demand. This paper proposes an intelligent inventory optimisation framework based on the Proximal Policy Optimisation (PPO) implemented with a deep reinforcement learning algorithm. We build a custom supply chain simulation environment with an eight-dimensional continuous state space to capture real-time inventory, demand features and seasonal patterns, alongside a continuous action space for precise ordering decisions and a business-calibrated reward function. Benchmark comparisons against Random, Fixed, EOQ, and MRP strategies show that the PPO achieves lower average total cost than MRP, though the performance gap is not statistically significant at α=0.05 with moderate Bayesian evidence. Nevertheless, the PPO model achieves stable convergence in training and strong scenario adaptability. Despite its higher bullwhip effect due to its demand responsiveness, the PPO’s effectiveness for intelligent inventory control is confirmed. Our framework offers a deployable model for industrial supply chain applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.