Skip to content
Conference

Optimization of last-mile delivery using dynamic path algorithms based on reinforcement learning

Jul 2026 · The 2026 International Conference on Optical Communication and Intelligent Algorithms (OCIA 2026) · Vol 14301, pp. 1430107 - 1430107-9 · 0 citations · 15 references
Engineering

TL;DR

The study uses a dynamic routing framework to create adaptive routing strategies, centered on a customized deep Q-network, that works well in variable traffic scenarios and effectively adapts to peak and off-peak delivery windows.

Abstract

Deep reinforcement learning has emerged as a transformative approach for solving the challenges of urban last-mile delivery, characterized by rapidly changing demand and unpredictable traffic conditions. The study uses a dynamic routing framework to create adaptive routing strategies, centered on a customized deep Q-network. By collecting real-time traffic data, vehicle status and delivery priority, the framework can formulate routing strategies. Using advanced data fusion algorithms, the method integrates heterogeneous sensors and traffic inputs. It also creates a composite reward function to balance operational costs, on-time delivery, and service quality. Simulations in a high-fidelity urban environment show that the proposed system can reduce the total delivery cost by 15 to 20 percent compared to static routing methods, while improving energy utilization and time efficiency. Sensitivity analysis shows that the method works well in variable traffic scenarios and effectively adapts to peak and off-peak delivery windows. These findings suggest that deep reinforcement learning and real-time data integration, along with multiobjective policy design, are effective solutions for optimizing modern urban logistics and last-mile delivery networks.

View source

Similar papers

Open access Aug 2026

Deep reinforcement learning-based adaptive traffic signal control in urban networks

This paper introduces a traffic signal control system using deep reinforcement learning to solve the congestion problems at signalized intersections under dynamic traffic conditions. The proposed framework is simulated with the help of MATLAB-SUMO co-simulation framework, the traffic signal control is modeled as a Markov Decision Process (MDP). State space includes traffic density, the queue length, vehicles waiting time, and the actual signal phase whereas the action space comprises of the possible selections of the signal phase. A Deep Q-Network (DQN) is utilized to estimate the optimal state-action value function so that green times can be dynamically allocated based on the changing traffic demand. Multi-objective reward functionality is based on the combined minimization of vehicle delay, queue length, and waiting time and maximization of traffic throughput. Experience replay and target network updates are used to stabilize the learning process. Simulation experiments are conducted in low, medium, and high traffic demand conditions to compare the work of the suggested framework with fixed-time control, actuated control, and classical tabular Q-learning methods. Experimental results demonstrate that the proposed framework reduces average delay by up to 33.9%, decreases queue length by 47.1%, and increases throughput by 25.5% compared to fixed-time control under high-demand conditions. Finally, scalability studies involving networks of up to 16 isolated signalized intersections were conducted to assess computational feasibility and robustness even when the size of the network grows. Comprehensively, the results prove the usefulness of deep reinforcement learning in the creation of intelligent and adaptive traffic signal control services in the city.

Manisha Aeri, K. Purohit, Lata Nautiyal et al. · 0 citations
Conference Aug 2026

A real-time dynamic vehicle path optimization framework for urban logistics based on deep reinforcement learning

This paper proposes a real-time dynamic vehicle path optimization framework for urban logistics, driven by advanced deep reinforcement learning techniques. The study describes the urban vehicle routing problem as a Markov decision process, integrating fleet operations, dynamic traffic conditions, and constantly arriving customer orders from heterogeneous realtime data streams. A graph-based neural network architecture for capturing complex spatio-temporal dependencies. This enables the system to learn and adapt to rapidly changing urban routing strategies. Both synthetic and real-world datasets are extensively tested. The proposed methods significantly reduce the operational cost and delivery latency. These methods are very effective compared to adaptive heuristic algorithms and traditional machine learning baselines. Maintaining robust performance under high traffic fluctuation and demand uncertainty is crucial for spatio-temporal feature extraction and network architecture optimization. The study demonstrates the feasibility and effectiveness of deep reinforcement learning technology in large-scale real-time logistics optimization in cities, and provides an important reference for the application of intelligent data-driven scheduling and path planning systems in complex urban networks.

Jinyan Wang, Hongjuan Cong · 0 citations
Open access Aug 2026

Application of Reinforcement Learning for Optimizing the Capacitated Vehicle Routing Problem

This study proposes a reinforcement learning (RL) framework for solving the deterministic single-depot Capacitated Vehicle Routing Problem (CVRP). The Capacitated Vehicle Routing Problem is formulated as a Markov Decision Process and a REINFORCE agent with a linear-softmax policy, incorporating Clarke-Wright savings features, is trained as a proof-of-concept prior to future deep architectures such as Deep Q-Network and Proximal Policy Optimization. The proposed agent is trained and evaluated on three reproducible synthetic benchmark datasets comprising 20, 50, and 100 customers, and its performance is compared with two conventional construction methods, namely Nearest Neighbor and Clarke-Wright Savings. The results show that the learned policy consistently converges to a stable routing strategy and outperforms the Nearest Neighbor heuristic on the medium- and large-scale instances, reducing total travel distance by 8.3% and 10.0% respectively, while remaining within 13.8-20.8% of the Clarke-Wright benchmark across all scenarios. Inference is completed within milliseconds once training is finished, indicating that the learned policy can be reused across new routing instances without restarting the optimization process. These findings demonstrate that reinforcement learning is a promising and scalable alternative to conventional heuristic and metaheuristic approaches for capacitated routing problems, particularly in dynamic logistics environments that require rapid and adaptive decision making.

Audrey Ariij Sya'imaa HS, Muhamad Rizky Aulia, Siti Hadiaty Yuningsih · 0 citations
Conference Jul 2026

A Sequential Constructive Reinforcement Learning Approach to Heterogeneous and Dynamic Multi-Vehicle Routing

Efficient vehicle routing is a fundamental problem in logistics and automated transportation systems, particularly in real-world settings where heterogeneous vehicle fleets operate under dynamic demand arrivals and stochastic travel conditions. Most existing methods are designed for static or homogeneous scenarios and struggle to scale when vehicle heterogeneity, evolving demands, and travel-time uncertainty must be jointly considered due to combinatorial complexity and inter-vehicle coupling. This paper proposes a lightweight deep reinforcement learning–based routing framework that reformulates multi-vehicle routing as a sequential decision-making process via a Sequential Route Construction strategy, enabling online re-optimization without high-dimensional joint action spaces. The proposed model jointly encodes road-network structure, global fleet-level states, and vehicle-specific attributes of the currently planning vehicle, allowing heterogeneous fleets to be coordinated under dynamic and stochastic environments. To ensure feasibility and training stability, feasibility-aware action masking and a Decision Sequence Equivalence (DSE) scheme are incorporated. Experimental results demonstrate that the proposed framework achieves competitive solution quality and high service fulfillment across heterogeneous and dynamic routing scenarios; compared with the strongest OR-Tools baseline in each setting, its inference-time search reduces objective values by 0.6–2.2% on HVRP and 0.8–1.7% on DSVRP while maintaining 0.991–1.000 fulfillment ratios with 0.495–2.019 s inference time.

Ming-Feng Li, Yao-Jiun Huang, Kuan-Han Chou et al. · 0 citations
Open access Jul 2026

Adaptive primal–dual Q-learning for electric vehicle route optimization on real-world charging networks

Electric Vehicles (EVs) are emerging as sustainable alternatives to internal combustion engine vehicles; however, efficient route planning remains a major challenge due to limited driving range, sparse charging infrastructure, and variable energy consumption patterns. Traditional shortest-path algorithms, such as Dijkstra’s and A*, often fail to account for EV-specific factors, including charging station availability, connector compatibility, and energy constraints. This study presents a comprehensive EV route optimization framework that integrates reinforcement learning (RL) with graph-based methods. A novel Dual Q–Adaptive Weighting model that balances reward and cost through a primal–dual learning mechanism is proposed. The framework learns energy-aware routing strategies from historical navigation experience. The model is compared against standard RL approaches—Q-Learning and Double Q-Learning—as well as enhanced variants of A* and Dijkstra’s algorithms that incorporate charging density and time-penalty considerations. Real-world EV charging infrastructure data from the Alternative Fuels Data Center (AFDC) and Placekey datasets are used to construct a clustered navigation graph via DBSCAN. Experimental results across multiple intercity routes show that the proposed Dual Q–Adaptive model achieves the highest route accuracy of 78.66%, outperforming Double Q-Learning (76.27%), Q-Learning (77.52%), and traditional A* (74.26%) and Dijkstra (60.92%) algorithms. A* and Dijkstra with modifications, use fewer charging stops than traditional algorithms. The Improvised algorithms provide substantial improvements over their baseline counterparts. The results demonstrate that reinforcement learning integrated with graph-theoretic optimization can enable scalable, infrastructure-aware, and efficient EV route planning.

Sarvesh Kumar, Rayappa David Amar Raj, Archana Pallakonda et al. · 0 citations
Open access Jul 2026

Dynamic traffic signal scheduling system based on adaptive quad agent Double Deep Q -network algorithm

Simulation results indicate that the proposed Adaptive Quad-Agent Double Deep Q-Network model effectively supports dynamic signal phase adaptation, minimizes congestion, and provides more accurate queue length estimations under complex traffic conditions.

Bharathi Ramesh Kumar, Sachin Salunkhe, S. Shinde et al. · 0 citations