Rare-Event Road-Traffic Fatality Prediction: A Reproducible Machine-Learning Benchmark with Time-Aware Validation, Calibration, and Interpretability
Urban road-safety agencies increasingly rely on administrative incident registries that contain many events but few fatalities. This study develops a reproducible machine-learning benchmark for rare-event road-traffic fatality prediction using an urban incident registry from Medellín, Colombia. The analysis is framed as a risk-ranking problem rather than as high-certainty binary classification, because fatal outcomes account for less than 1% of the records. The benchmark compares a prevalence-only reference, logistic regression, CART, Random Forest, and XGBoost under a shared preprocessing and time-aware validation design. Historical records are used for training and later observations are held out for testing, reducing temporal leakage and approximating prospective use. Model performance is evaluated with metrics suited to severe class imbalance, including ROC-AUC, PR-AUC, Youden-based threshold summaries, Precision@1%, bootstrap uncertainty intervals, calibration diagnostics, feature-importance analysis, and sensitivity checks. The results show that the available registry variables support risk enrichment but not high-precision fatality classification. The study contributes a transparent baseline for computational road-safety research and clarifies the limits of registry-based prediction when exposure, traffic-flow, roadway, infrastructure, weather, and post-crash response variables are not available.