Skip to content

A robust ensemble learning approach for human-factor based accident severity prediction: Insights from a decade of Brazilian highway data

Jul 2026 · Intelligent Data Analysis · 0 citations · 25 references

TL;DR

A robust intelligent framework to classify truck driver fatalities on Brazilian federal highways by specifically isolating human-factor variables is proposed, and SHAP (SHapley Additive exPlanations) is integrated to provide a knowledge-based interpretation of the model's decisions.

Abstract

Traffic accidents involving heavy vehicles remain a critical safety challenge, with human factors being the primary contributors to fatality risks. While machine learning has been widely used for accident prediction, the lack of model interpretability often hinders its application in real-world policy-making. This study proposes a robust intelligent framework to classify truck driver fatalities on Brazilian federal highways by specifically isolating human-factor variables. Leveraging an extensive dataset from the Brazilian Federal Highway Police (2013–2025), we conducted a comparative analysis of three ensemble-based algorithms: Random Forest, XGBoost, and LightGBM. To ensure model stability and generalization, hyperparameter optimization was executed using RandomizedSearchCV with five-fold cross-validation. The experimental results demonstrate that while Random Forest achieved high training accuracy, LightGBM emerged as the superior model for safety-critical deployment, achieving a balanced ROC-AUC of 0.843 and a superior recall, effectively minimizing life-threatening false negatives. Furthermore, this research integrates SHAP (SHapley Additive exPlanations) to provide a knowledge-based interpretation of the model's decisions. The XAI analysis reveals that “Accident Type” and specific “Human Factor Categories” are the most significant predictors of fatality. The findings provide a transparent, data-driven decision support tool for transportation authorities to implement targeted interventions, bridging the gap between complex black-box models and actionable road safety knowledge.

View source

Similar papers

Open access Jul 2026

Human Error Analysis in Maritime Accidents: A Hybrid Machine Learning Approach for Enhanced Predictive Modeling

This study presents a novel hybrid machine learning approach for analyzing human error factors in maritime accidents using HELCOM accident data. We developed an advanced classification model that combines Random Forest, Gradient Boosting, XGBoost, LightGBM, and neural networks through a soft voting ensemble mechanism. Our methodology includes enhanced data preprocessing techniques, sophisticated feature engineering, and class imbalance correction through SMOTE resampling. The visualizations produced meet publication-quality standards with optimized color schemes, typography, and statistical representations suitable for high-impact journals. The hybrid model achieved 84% accuracy with macro-averaged F1-score of 0.58 across eight accident classes, identifying key human factors contributing to maritime incidents. Our findings indicate that specific human element factors have significant correlations with particular accident types, offering valuable insights for maritime safety policy development and accident prevention strategies. This research contributes to the growing field of data-driven maritime risk assessment by providing a robust methodological framework for human error analysis in the maritime domain.

Prabhat Nigam, Neeraj Anand, Manan Bhasin et al. · 0 citations
Open access Aug 2026

Traffic Accident Severity Prediction Based on Multi-Model Comparison

Traffic accident severity prediction is important for improving traffic safety management and supporting risk assessment and prevention. However, current studies have several limitations with respect to the systematic approach for comparing models and the comprehensiveness of evaluation metrics. This paper aims to systematically assess the overall performance of various machine learning models on a unified experimental framework to predict the severity of accidents in a binary classification task. This study randomly sampled 500,000 records from the US Accidents dataset and used them as the experimental sample. Following data cleaning, missing value handling, feature engineering, and categorical variable encoding, a comparison of four representative models was performed: Logistic Regression (LR), Random Forest (RF), XGBoost and Multi-Layer Perceptron (MLP). With the imbalanced nature of the data, this paper evaluates model performance using accuracy, AUC, ROC curves, and confusion matrices to assess the overall classification performance of different models and their ability to identify the severe accident class. The results show that the tree-based ensemble models Random Forest and XGBoost outperform Logistic Regression and MLP in terms of overall predictive performance, with XGBoost exhibiting better overall performance. Variables that contribute significantly to the model’s predictions include traffic control facilities, time factors, spatial location, and accident-affected distance. The results indicate that machine learning techniques could be useful in traffic accident severity prediction and could be a reference for traffic safety risk assessment. Meanwhile, class imbalance, feature interpretability and model generalizability are challenges that need further investigation and resolution.

Zi-Yu Cao · 0 citations
Open access Jul 2026

Rare-Event Road-Traffic Fatality Prediction: A Reproducible Machine-Learning Benchmark with Time-Aware Validation, Calibration, and Interpretability

Urban road-safety agencies increasingly rely on administrative incident registries that contain many events but few fatalities. This study develops a reproducible machine-learning benchmark for rare-event road-traffic fatality prediction using an urban incident registry from Medellín, Colombia. The analysis is framed as a risk-ranking problem rather than as high-certainty binary classification, because fatal outcomes account for less than 1% of the records. The benchmark compares a prevalence-only reference, logistic regression, CART, Random Forest, and XGBoost under a shared preprocessing and time-aware validation design. Historical records are used for training and later observations are held out for testing, reducing temporal leakage and approximating prospective use. Model performance is evaluated with metrics suited to severe class imbalance, including ROC-AUC, PR-AUC, Youden-based threshold summaries, Precision@1%, bootstrap uncertainty intervals, calibration diagnostics, feature-importance analysis, and sensitivity checks. The results show that the available registry variables support risk enrichment but not high-precision fatality classification. The study contributes a transparent baseline for computational road-safety research and clarifies the limits of registry-based prediction when exposure, traffic-flow, roadway, infrastructure, weather, and post-crash response variables are not available.

Erika María López-López, O. Bru-Cordero, C. D. Correa-Álvarez · 0 citations
Open access 2026

A Statistical and machine learning framework for analyzing road traffic crash severity in Sri Lanka

This study analyses 397,850 police-reported crashes in Sri Lanka from 2010 to 2020 using a multinomial logit model, validated through Random Forest and XGBoost, to identify key factors influencing crash severity across four outcome levels. Results reveal four findings not apparent from conventional crash statistics. First, motorcycle–vulnerable road user crashes (OR = 202, 10.24% of crashes) produced higher fatal odds than heavy vehicle–motorcycle interactions (OR = 143, 6.86%), despite the latter receiving greater policy attention. This suggests that motorcycle-related risk to vulnerable road users is considerably underappreciated in current safety programs. Second, single-vehicle motorcycle crashes (OR = 46.2, 9.05%) were three times more fatal per incident than motorcycle–passenger vehicle collisions (OR = 14.6), implicating loss of control and emergency response delays as under-recognized contributors. Third, passengers falling from buses yielded a fatal odds ratio of 2,337, the highest in the model, despite representing only 1.50% of crashes, indicating a concentrated fatality risk understated in national reporting. Fourth, dusk and dawn conditions (OR = 2.24, 17.27% of crashes) were found to be equally as dangerous as completely unlit nighttime roads (OR = 2.26), while fog and mist produced the highest lighting-related fatal odds ratio of 4.71. The findings indicate that targeted interventions addressing motorcycle–vulnerable road user interactions, bus passenger safety, and transitional lighting conditions offer the greatest potential for reducing fatal crash outcomes in Sri Lanka.

D. Bhagya, N. Jayantha, H. Pasindu · 0 citations
Open access Jul 2026

Data-driven road safety enhancement: neural network–based accident classification and safe route identification using spatial network analysis

Road traffic accidents continue to pose a significant challenge to public safety, resulting in substantial human suffering, economic losses, and increasing pressure on transportation systems. This study proposes a data-driven intelligent transportation framework that integrates machine learning and spatial network analysis to support accident severity prediction, risk-aware route recommendation, and emergency response. A comprehensive dataset comprising traffic conditions, weather information, temporal attributes, roadway characteristics, vehicle information, and driver-related factors was analysed to identify the key determinants of accident severity. Multiple machine-learning models were evaluated, and the Multi-Layer Perceptron (MLP) classifier achieved the highest predictive performance, attaining an overall accuracy of 91.2%. To transform predictive outcomes into practical safety interventions, the proposed framework combines accident severity prediction with GPS-enabled spatial network analysis to identify high-risk road segments and recommend safer alternative routes. In addition, an automated SMS notification mechanism is incorporated to provide location-aware emergency alerts when high-risk situations are detected. The integration of predictive analytics, spatial risk assessment, safety-oriented routing, and emergency communication establishes a comprehensive decision-support framework for proactive accident prevention and transportation-safety management. The experimental results demonstrate that the proposed approach can effectively support safer mobility, improved situational awareness, and enhanced emergency response within intelligent transportation environments.

V. Naresh, Ayyappa Dullam · 0 citations