A reproducible Deep Reinforcement Learning baseline for tourist bus route planning: a case study in Ha Noi
Abstract
Planning itineraries for urban tourist buses involves a critical trade-off between maximising attraction coverage and strictly adhering to operational constraints. While recent literature has introduced highly complex optimisation models, there remains a notable lack of a transparent, reproducible baseline tailored specifically for this domain. To address this gap, this paper presents a straightforward Deep Reinforcement Learning (DRL) baseline framework that pairs a Double Deep Q-Network (Double DQN) with a deterministic pathfinding module. Operating solely on fundamental static inputs, such as a bus-feasible road graph, points of interest (POIs), and fixed time budgets, the model generates structurally feasible routes within a simulated environment without requiring complex dynamic data. The framework is evaluated through a simulation-based case study in Ha Noi City, simulating both morning and afternoon shifts under a four-hour service constraint. Empirical results indicate that the Double DQN architecture exhibits improved training stability, yielding ~13% improvement in POI coverage compared to a Standard DQN baseline across different times of day. Also, the proposed model achieves ~90% time-budget utilisation rate, improving upon ~79% observed in standard methods, while converging roughly twice as fast. These findings provide a simulation-based DRL baseline for future comparative studies in intelligent tourism mobility.