Skip to content
Open access

Explainable machine learning framework for foodborne disease outbreak prediction in Eastern Province, Saudi Arabia: case study

Aug 2026 · Frontiers in Public Health · 0 citations · 23 references

TL;DR

A machine learning-based model for predicting outbreaks, explainability, and spatial risk propagation, validated through a multiyear data set of an epidemiological nature from 12 cities in the Eastern Province of Saudi Arabia (2021–2025).

Abstract

The occurrence of foodborne diseases is a considerable public health issue, especially in areas that are quickly becoming urbanized with intricate food delivery systems. In this paper, we present a machine learning-based model for predicting outbreaks, explainability, and spatial risk propagation, validated through a multiyear data set of an epidemiological nature from 12 cities in the Eastern Province of Saudi Arabia (2021–2025). The final data set includes 61 cases and 13 engineered features. In the current research, the proposed architecture uses XGBoost to predict outbreaks, alongside using the random forest for predicting severity and support vector machine (SVM) for comparisons. The XGBoost classifier demonstrates an evenly balanced performance (accuracy = 0.85, precision = 0.78, recall = 0.78) on the testing set. Due to the size of the dataset, the results are provided with the estimation of uncertainty (rather than the exact numbers). Using leakage-safe repeated stratified cross-validation, the mean AUC equals 0.64 [95% interval = (0.20, 1.00)], and the leave-one-year-out validation method is not stable (mean AUC 0.47). Differences between the models (e.g., better single-split cross-validation AUC for SVM) are within confidence intervals. Interpretability is improved by using the SHAP framework to measure feature importance, which shows that the main factors are hospitalization and the severity of symptoms. The graph module also helps in understanding the propagation of disease risk between cities, highlighting the importance of well-connected metropolitan areas as disease hubs. Moreover, the use of a locally deployed Mistral LLM makes the generated explanations more readable. The findings show that our approach presents an appropriate balance of predictiveness, interpretability, and spatial knowledge. With only 61 data points and 11 outbreaks reported, this research is clearly not meant to be an early warning system, but rather a proof-of-concept on how one might be designed. In order to ensure reproducibility, the preprocessing pipeline and synthetic dataset generator have been made available to the community.

Read PDF

Similar papers

Open access Jul 2026

Use of Geospatial Big Data Intelligence Software for Epidemic Forecasting and Public Health Decision-Making:

The development and assessment of a machine learning-driven early warning system for infectious disease prediction using geospatial big data from South-Western Nigeria, at the level of Local Government Area outperform conventional surveillance systems in developing countries.

I. Adewumi, N. Bakare, W. Ajayi et al. · 0 citations
Review Open access 2023

Computational Intelligence Models for Disease Outbreak Prediction

A comprehensive review of CI models for outbreak prediction, comparing supervised and unsupervised methods such as Support Vector Machines (SVM), Random Forest (RF), Artificial Neural Networks (ANN), Long Short-Term Memory (LSTM), and hybrid models.

Z. Abdullahi · 0 citations
Open access Aug 2026

A Machine Learning-Based Health Issues Detection System in Crude Oil Exploitation Host Communities in Ondo State

Crude oil exploitation in Ondo State, Nigeria, causes severe environmental pollution and adverse health outcomes like respiratory disorders and water-borne illnesses. Current surveillance systems remain reactive and inefficient at detecting early warning signs. To address this gap, this study developed a machine learning-based health issues detection system (HIDS) to enhance early detection in vulnerable populations. Due to data collection constraints, synthetic dataset of 5000 records was generated based on epidemiological patterns and environmental indicators including air pollution, water contamination, and proximity to gas flare sites. The Random Forest algorithm was selected for its robustness with non-linear relationships. Evaluated using an 80:20 split, the model achieved strong performance: accuracy 82.34%, precision 83.12%, recall 81.45%, specificity 83.56%, F1-score 82.23%, and ROC-AUC 85.23%. Feature imporatnce analysis identified environmental factors as dominant predictors. Air pollution index (0.287), water contamination index (0.245), and distance to gas flare sites (0.198) collectively accounted for 73% of predictive power. Additionally, a three-tier risk stratification framework was established where low (0.00-0.40), moderate (0.41-0.70), and high (0.71-1.00) to facilitate proactive healthcare interventions and efficient resource allocation. This paper confirms the feasibility of applying machine learning for public health surveillance in resource constrained environments. It provides policymakers and healthcare providers with vital tool for mitigating health risk in crude oil host communities through data-driven early warning systems rather than reactive approaches, ultimately improving health outcomes in environmentally vulnerable populations affected by extractive industry activities.

Olutomisin M. Orogbemi, Balogun John Tope · 0 citations
Open access Aug 2026

An Explainable Heterogeneous Stacking Ensemble Framework for Fish Health Prediction in Intelligent Aquaculture

Aquaculture is a vital element of global food security and economic sustainability. The changing environment from time to time and the rapid spread of diseases in aquaculture leading to huge economic losses, still pose a major challenge in the maintenance of healthy fish. Timely intervention for rapid and accurate fish health forecasting and diagnosis will decrease the fish mortality and increase the efficiency of fish production. In the current paper, a stacking model using heterogenous ensemble learning will be developed using three machine learning algorithms including the Random Forest algorithm, Support Vector Machine (SVM) and XGBoost together with Logistic Regression as the meta-classifier to predict fish health. Data preprocessing stage should be conducted prior to developing a model which will include missing value treatment, feature scaling, feature encoding and irrelevant feature elimination. The efficiency and generality of the proposed model will be estimated via using stratified 10-fold cross-validation technique. The quality of the model will be estimated by using the criteria including accuracy, precision, recall, F1 Score, confusion matrix and ROC curve analysis. In addition, SHapley Additive exPlanations will be applied to increase the explainability of the proposed model by showing the contribution of each environmental and biological factor into predicting fish health. The experimental results reveal that the heterogeneous stacking model can outperform the individual base learners and provide a transparent decision explanation. The framework is proposed to be utilized for intelligent fish health monitoring and contributing to sustainable aquaculture through predicting diseases in time and managing the aquaculture farms effectively, interpretable and scalable way.

Ranjan Kumar, Savita Choudhary · 0 citations
Open access 2026

Case-Based Reasoning Model for Predicting the Malaria Cases

Beyond predictive accuracy, qualitative criteria including explain-ability, transparency, and adaptability were evaluated, further highlighting the superiority of CBR 2 over conventional black-box models.

Konan N’gatta Aimé Kouassi, Koffi Kouakou Ive Arsene, Goore Bi Tra · 0 citations