Aug 2026· Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2· 0 citations· 36 references
TL;DR
CityWeave is proposed, a VLM-based framework for urban D2D mobility planning that integrates the Who--When--Where--How (3W1H) reasoning paradigm with a two-stage training scheme and introduces a unified User--World Grounding (UWG) module that enforces navigation-based world constraints and evaluates personalization with respect to the user profile.
Abstract
Urban door-to-door (D2D) mobility planning is a core task for AI-powered smart cities, requiring models to capture individual mobility behavior and generate optimized plans under real-world urban constraints such as network connectivity and service schedules. Existing methods face fundamental limitations. Optimization-based approaches rely on static costs and fail to capture individual-specific preferences. LLM-agent-based approaches often have weak spatio-temporal reasoning and unstable constraint tracking, which reduces feasibility and reliability. In this study, we propose CityWeave, a VLM-based framework for urban D2D mobility planning that integrates the Who--When--Where--How (3W1H) reasoning paradigm with a two-stage training scheme. CityWeave learns this paradigm through supervised fine-tuning and is further improved by reinforcement learning based enhancement. A dataset of 180,000 real-world samples from 80,000 users is constructed to support training and evaluation. The model learns to identify user needs (Who), reason over departure and arrival time windows (When), read maps and spatial topology (Where), and invoke routing tools (How) to generate feasible plans. We further introduce a unified User--World Grounding (UWG) module that enforces navigation-based world constraints and evaluates personalization with respect to the user profile. Extensive experiments show that CityWeave achieves a state-of-the-art Final Pass Rate of 64.7% and a Commonsense Pass Rate of 92.4%, outperforming both conventional non-LLM planning pipelines and strong LLM-agent baselines. These results demonstrate that structured reasoning over human mobility behavior, combined with explicit user and world grounding, offers a practical path toward reliable and personalized planning agents for smart urban transportation systems.
Urban planning is a real-world spatial optimization problem that requires selecting feasible actions from large candidate spaces under practical objectives such as cost and service quality. Existing optimization and reinforcement learning methods are effective for fixed formulations, but often depend on task-specific r...
Wen-Tao Zhang, Jing-Yuan Wang, Ze-Tong Zhou et al.· 0 citations
Individual mobility trajectories support urban analysis and location-based services, yet most trajectory generators require observations from their deployment city. This assumption excludes precisely the cities where trajectories are unavailable even though points of interest (POIs) and their attributes can be obtained...
Yi-Di Wang, Yun-He Zhang, Bang-Chao Deng et al.· 0 citations
Human mobility generation aims to synthesize realistic point-of-interest (POI) visitation trajectories and has become an important tool for travel behavior modeling, transportation management, and urban planning. Existing diffusion-based methods achieve high fidelity but require per-city generation, given the inherent...
Zhou-Fu Wang, Bao-Shen Guo, Zhi-Qing Hong et al.· 0 citations
Experimental results show that the proposed P-LLM framework outperforms traditional scheduling methods and existing reinforcement learning baselines, and maintains stable and consistent performance across different real-world scenarios, time periods, fleet sizes, and order volumes.
Jia-Xin Tan, Xiao-Hui Huang, Nan Jiang et al.· Applied intelligence (Boston...· 0 citations
This work presents NeuralParker, a reinforcement learning-based hybrid planner for arbitrary-pose parking that encodes full-environment obstacle and boundary geometry in a target-relative vertex representation, allowing the policy to retain route-defining context throughout the approach.
Zihan Wang, Baixiang Huang, Yang Guan et al.· 1 citation
Vision-and-language navigation (VLN) has advanced rapidly in static indoor environments, but robots operating in human-populated spaces must ground language while responding to moving pedestrians and social-safety constraints. We present DPed-VLN, a Habitat 3.0 benchmark for dynamic-pedestrian VLN that couples 33,093 n...
Hao-Jie Dai, Xiang-Yi Wang, Liu-Yi Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.