Skip to content
Open access

A Reinforcement Learning Framework for Traveling Salesman and Vehicle Routing Problem with Drones

Aug 2026 · Drones · Vol 10, pp. 616 · 0 citations · 31 references

TL;DR

A reinforcement learning method with a shared attention encoder and a hierarchical dual-decoder architecture, where truck–drone coordination is achieved by first decoding the truck’s next node and then conditionally decoding the drone action.

Abstract

The Traveling Salesman Problem (TSP) and the Vehicle Routing Problem (VRP) are two classical combinatorial optimization problems. In recent years, their drone-assisted variants, the Traveling Salesman Problem with Drones (TSP-D) and the Vehicle Routing Problem with Drones (VRP-D) have attracted growing attention. Generally, these problems are solved using exact algorithms or metaheuristic algorithms. However, as the problem complexity increases and the scale of instances grows, these approaches often become less efficient. In this paper, we propose a reinforcement learning method with a shared attention encoder and a hierarchical dual-decoder architecture, where truck–drone coordination is achieved by first decoding the truck’s next node and then conditionally decoding the drone action. To further explore the solution space of large-scale instances, the proposed method adopts a multi-rollout learning strategy. We conducted experiments on large-scale TSP-D and VRP-D instances, and the results show that this model outperforms traditional metaheuristic algorithms in terms of both solution quality and computational efficiency.

Read PDF

Similar papers

#reinforcement learning Open access Sep 2026

Reinforcement learning for initializing genetic algorithms in vehicle routing

This work introduces an optimization framework where a reinforcement learning agent is trained on prior instances and quickly generates initial solutions, which are then further optimized by a genetic algorithm, enabling real-time and interactive routing at scale.

Ido Greenberg, P. Sielski, Hugo Linsenmaier et al. · 1 citation

for Vehicle Routing Problems

Deep Policy Dynamic Programming is proposed, which aims to combine the strengths of learned neural heuristics with those of DP algorithms, and prioritizes and restricts the DP state space using a policy derived from a deep neural network, which is trained to predict edges from example solutions.

W. Kool, H. van Hoof, J. Gromicho et al. · 0 citations
Open access Aug 2026

A solution method for the traveling salesman problem based on multi-scale features and dynamic optimization

A multi-scale deep optimization model based on an encoder-decoder architecture that validates the effectiveness of the multi-scale EMA and Triplet-Reasoning mechanisms, providing a new direction for deep learning-based graph optimization research.

Yu-Ting Xie, Qianqian Duan · 0 citations
Review Open access Sep 2026

Advanced Optimization Methods for the Knapsack, Traveling Salesman, and Close-Enough Traveling Salesman Problems: A Survey and Case Studies

Combinatorial optimization problems (COPs), including the Knapsack Problem (KP), the Traveling Salesman Problem (TSP), and regional variants such as the Close-Enough Traveling Salesman Problem (CETSP), constitute fundamental models for addressing complex decision-making tasks in modern computational systems. Their comp...

S. El Kafhali, Mohamed Abid, Mohamed Hanini · 1 citation
#artificial intelligence Preprint Sep 2026

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

The experimental results demonstrate that RLEA outperforms the previous state-of-the-ar method, achieving a 16.67% higher success rate while significantly reducing runtime errors, and validate that integrating reinforcement learning with LLM-based reasoning is highly effective for automated optimization modeling.

Yi Chen, Zi-Pei Yu, Jia-Hai Wang et al. · 1 citation
Open access 2026

An Adaptive Large Neighborhood Search for the Multiple Traveling Salesman Problem With Backup Coverage

The Multiple Traveling Salesman Problem with Backup Coverage (mTSP-BC) is a vehicle routing variant in which all vehicles must remain within a maximum pairwise distance at every instant during their traversal, imposing spatiotemporal interdependence among routes. This constraint models real-world scenarios such as mili...

Jonathan Cardozo Maciel, Guilherme Dhein, O. B. D. de Araújo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.