Machine-Learning-Based Traffic State Prediction in Car–Bicycle Mixed Traffic Using Synthetic Data
This study explores the use of machine learning models to predict traffic conditions in mixed car–bicycle traffic environments. A synthetic dataset was developed from numerical evaluations of traffic flow theory, capturing a wide range of multimodal traffic scenarios. Random forest (RF), multi-layer perceptron (MLP), and linear regression models were trained to estimate key traffic metrics, including output flow, delay, and density. The analysis focuses on model performance under different data splits, especially when sorting by variables such as initial car flow and bicycle flow. Results show that, while RF performs well for previously observed traffic conditions, MLP offers stronger generalization to unseen traffic conditions, particularly in high-flow and high-density regimes. However, prediction performance varies depending on the input variable used for sorting and the distribution of training data. These findings underscore the importance of balanced, diverse datasets and support the use of data-driven models for traffic state estimation in multimodal urban networks.