A Bayesian spatiotemporal hierarchical model designed specifically for inductive forecasting that is able to accurately estimate O-D flows and provide robust inductive forecasts with full uncertainty quantification, which is essential for robust decision making in downstream applications, such as stochastic network optimization and facility location problems.
Abstract
Accurate forecasting of origin-destination (O-D) demand is critical for transportation network design, planning, and operational management. An underexplored challenge is inductive forecasting: predicting O-D flows for new or unobserved locations, which is essential for evaluating network expansions and new facility placements. However, most existing approaches either are transductive (limited to fixed observed networks) or are deep learning models that lack interpretability, fail to provide uncertainty quantification, and often overlook critical data constraints such as structural zeros. To this end, we develop a Bayesian spatiotemporal hierarchical model for probabilistic O-D demand estimation, designed specifically for inductive forecasting. We assume total trip generation is given and model the destination-choice counts using a multinomial distribution. To handle the high-dimensional probability tensor, which features both simplex and structural zero constraints, we parameterize it via a masked-centered softmax transformation of a latent utility tensor. We then model this latent utility tensor using a CANDECOMP/PARAFAC (CP) tensor factorization to parsimoniously capture spatiotemporal patterns. Crucially, we impose Gaussian process (GP) priors on the spatial and temporal latent factors; the GP over the network provides the principled statistical mechanism to make inductive predictions for new locations. For posterior inference, we propose an Markov chain Monte Carlo algorithm. We validate the proposed model on synthetic data and two real-world O-D data sets. Results confirm our model’s ability to accurately estimate O-D flows and provide robust inductive forecasts with full uncertainty quantification, which is essential for robust decision making in downstream applications, such as stochastic network optimization and facility location problems.
Funding: This research is supported by the Natural Sciences and Engineering Research Council (NSERC) of Canada [Discovery Grant RGPIN-2025-04479].
Supplemental Material: The online appendix is available at https://doi.org/10.1287/trsc.2026.0063 .
Entry-only automatic fare collection systems record boardings but not alightings, preventing direct construction of origin-destination (OD) matrices. This study develops a Hierarchical Bayesian Latent-Destination (HBLD) model that combines station-hour boarding and inferred alighting demand with passenger card histories. Trip-chain destinations are treated as noisy evidence with a reliability parameter, allowing destination uncertainty to propagate into OD flows. The model was applied to 838,305 bus tap-ins collected in Changzhou in May 2025 and linked to stop-network and hourly weather data. It estimates destination distributions over feasible downstream and reverse-direction through-terminal stops using network, time-of-day, weather, and smoothed historical demand effects. A Bayesian personalization layer uses prior card trips and reverts to the shared trip-level distribution when history is unavailable. Fitted by stochastic variational inference and evaluated on the final week, HBLD outperformed the strongest baseline. Observed boarding patterns consistently improved prediction, especially without card history, while inferred alighting patterns helped only when trip-chain evidence was strongly trusted. The model captured travel consistent with through-terminal riding and bus-assisted road crossing and estimated destinations for trips unresolved by deterministic chaining. Because true alightings were unavailable, scores measure agreement with trip-chain outputs rather than actual destination accuracy. HBLD provides uncertainty-aware destination predictions and OD matrices for service management, planning, scheduling, and resource allocation.
A Tensor Network Extended Kalman Filter (TNEKF) framework for short-term metro OD demand forecasting that consistently outperforms ARIMA, conventional EKF, and several state-of-the-art spatiotemporal prediction models in terms of MAE, RMSE, and MAPE.
Aijing Su, Bing Wu, Xiaoxing Fang· ISPRS International Journal...· 0 citations
An integrated prediction-and-visualisation pipeline that transforms complex data distributions into actionable visual analytics, such as interpretable station-to-station demand heatmaps via interactive GIS Folium layers is implemented, providing an operationally robust framework to support smart-city transportation management and build more sustainable urban transit systems.
Berna Çalışkan· Journal of Data Analytics an...· 0 citations
Accurate passenger flow prediction is crucial to the development of efficient public transit systems, enabling optimal resource allocation, dynamic route planning, and robust congestion mitigation. However, forecasting highly sparse origin-destination (OD) matrices remains a persistent challenge due to extreme data overdispersion and the “zero convergence” problem, where models gravitate toward predicting zeros to minimize global error at the expense of edge-level accuracy. Traditional sequential forecasting models often exacerbate these issues by relying on extensive historical data sequences, which compounds memory overhead and limits real-time scalability. In order to address these limitations, this study proposes a novel, compact, physics-informed generative framework (the Dual-Head VAE-GMM) capable of forecasting transit demands using only a single historical time step as input. The proposed architecture incorporates a Gaussian Mixture Model (GMM) latent prior to capture multi-modal mobility regimes and utilizes a dual-head decoder that explicitly decouples binary network topology edge prediction from continuous volume regression. To ensure predictions align with real-world flow dynamics, the model is governed by a composite objective function that integrates physics-informed graph spectral regularizers and domain-aware mass conservation laws into a generalized variational lower bound. Evaluations on a comprehensive smart card dataset from the 748-station Seoul metropolitan subway network indicate that the proposed approach successfully mitigates zero-inflation and performs significantly better than traditional parametric and deep-learning baselines. The model achieves a mean absolute error of under two passengers with an inference latency of about 57 milliseconds for a 60-minute forecasting horizon.
Mark Mpabulungi, Changhun Kang, Keemin Sohn et al.· IEEE Access· 0 citations
Demand-responsive transit (DRT) has emerged as a flexible mobility solution for addressing service blind spots in areas underserved by conventional public transportation. However, empirical research on DRT demand estimation remains limited, especially in the context of regional heterogeneity and zero-inflated demand. This study proposes a data-driven DRT demand estimation framework using operational records from two contrasting regions in Incheon, South Korea. Yeongjongdo is a tourism- and airport-oriented area with a high floating population and strong temporal variability, while Geomdan New Town is a residential district with relatively stable travel patterns. A grid-based origin–destination (O–D) modeling structure was employed, incorporating spatial variables such as land use, population dynamics, facility distribution, and public transport accessibility. Multiple region-specific machine learning models were developed and evaluated to estimate planning-oriented mean daily demand for each O–D-hour. A comparative analysis of alternative model configurations showed that the appropriate model structure varied according to regional demand conditions. In Yeongjongdo, the proposed two-stage model, which combines demand-occurrence classification and conditional regression, achieved the best overall performance, with a test MAE of 0.0022 and RMSE of 0.0105. These values represented reductions of 37.1% and 42.0%, respectively, relative to the strongest direct regression benchmarks. In contrast, direct regression was more effective in Geomdan New Town, achieving a test MAE of 0.0071 and RMSE of 0.0190. These results indicate that the relative performance of direct regression and two-stage prediction may vary across regional demand contexts.
Yunji Jang, Eun Hak Lee, Jiho Yeo et al.· Scientific Reports· 0 citations