Skip to content

Reinforcement learning for initializing genetic algorithms in vehicle routing

Sep 2026 · Communications AI & Computing · Vol 1 · 1 citation · 69 references
Vehicle Routing Optimization Methods

TL;DR

This work introduces an optimization framework where a reinforcement learning agent is trained on prior instances and quickly generates initial solutions, which are then further optimized by a genetic algorithm, enabling real-time and interactive routing at scale.

Abstract

Vehicle routing problems (VRP) are an extension of the Traveling Salesperson Problem and are a fundamental NP-hard challenge in combinatorial optimization. Solving VRP in real-time at large scale has become critical in numerous applications, from growing markets like last-mile delivery to emerging use-cases like interactive logistics planning. Such applications involve solving similar VRP instances repeatedly, yet current state-of-the-art solvers treat each instance on its own without leveraging previous examples. We introduce an optimization framework where a reinforcement learning agent is trained on prior instances and quickly generates initial solutions, which are then further optimized by a genetic algorithm. This framework, Evolutionary Algorithm with Reinforcement Learning Initialization (EARLI), consistently outperforms current state-of-the-art solvers under limited time budgets. For example, EARLI handles vehicle routing with 500 locations within one second, 10x faster than current solvers for the same solution quality, enabling real-time and interactive routing at scale. EARLI can generalize to new data, as demonstrated on real e-commerce delivery data of a previously unseen city.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

The experimental results demonstrate that RLEA outperforms the previous state-of-the-ar method, achieving a 16.67% higher success rate while significantly reducing runtime errors, and validate that integrating reinforcement learning with LLM-based reasoning is highly effective for automated optimization modeling.

Yi Chen, Zi-Pei Yu, Jia-Hai Wang et al. · 1 citation
Open access Aug 2026

Application of Reinforcement Learning for Optimizing the Capacitated Vehicle Routing Problem

These findings demonstrate that reinforcement learning is a promising and scalable alternative to conventional heuristic and metaheuristic approaches for capacitated routing problems, particularly in dynamic logistics environments that require rapid and adaptive decision making.

Audrey Ariij Sya'imaa.HS, Hilda Azkiyah, Khandker Farid Uddin Ahmed · 0 citations

for Vehicle Routing Problems

Deep Policy Dynamic Programming is proposed, which aims to combine the strengths of learned neural heuristics with those of DP algorithms, and prioritizes and restricts the DP state space using a policy derived from a deep neural network, which is trained to predict edges from example solutions.

W. Kool, H. van Hoof, J. Gromicho et al. · 0 citations
Conference Open access 2026

LLM-enhanced Dynamic Fleet Planning with Hierarchical Multi-agent Reinforcement Learning Framework

The proposed hierarchical multi-agent proximal policy optimization framework can reduce total airlines' operational costs—including direct operating cost and capital cost and achieves a computation speedup in comparison with a conventional optimization baseline.

Li-Jing Liu, James M. Shihua, Qi-Yu Yan et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 7, 2026

Discovering the value of humanistic inquiry

Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.

Microsoft Research Blog Oct 7, 2026

Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses

Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.