The Evaluation Gap: Why Standard Accuracy Metrics and Drift Detectors Understate the Deployment Risk of Crash-Severity Prediction Models
Road authorities increasingly use machine-learning models to rank high-risk locations and to allocate scarce safety resources. These models are almost always judged by their predictive accuracy on a random hold-out sample. This study shows that practice hides an important risk. We evaluate crash-severity models the way...