A Clustering-Based Descriptive Analysis of Urban Road Traffic Accidents
Road traffic accidents are a leading cause of death and disability worldwide, and their burden falls disproportionately on low- and middle-income countries, yet data-driven analyses remain scarce for Sudan. This study presents a descriptive and unsupervised-learning analysis of 2,218 road traffic accidents recorded across the localities of Khartoum State. After cleaning, translating, and feature-engineering the records, an exploratory analysis is conducted through a series of descriptive visualizations spanning the spatial, temporal, vehicular, demographic, and severity dimensions of the data. K-means clustering is then applied within a principal-component space and validated using the elbow method, silhouette analysis, the Calinski–Harabasz and Davies–Bouldin indices, and a hierarchical-clustering cross-check, to derive latent accident typologies. A five-cluster solution separates a dominant high-volume low-severity pedestrian and motorcycle stratum from a smaller but critical high-fatality cluster, a bus-and-collision cluster, and an outlier group, with fatality rates ranging from 1.0% to 99.0%. The analysis reveals that the pedestrian and the afternoon/early evening hours account for almost half of all accidents, and that an interpretable risk segmentation can be derived, with direct consequences for the road-safety policy in Khartoum.