Skip to content

Leveraging machine learning for symptom-based malaria diagnosis: a predictive model approach

Jun 2026 · Neural computing & applications (Print) · Vol 38 · 0 citations · 60 references
Computer Science

TL;DR

A region-specific malaria prediction model for clinical diagnosis that uses patient demographics and clinical symptoms has the potential to assist in the development of a clinically based malaria diagnostic system.

View source

Similar papers

Open access Jul 2026

Explainable AI for Malaria Diagnosis: Comparative Analysis of ML Models Using Random Forest Feature Selection and SHAP Interpretability

This framework illustrates, without any claim of clinical validity, how a leakage-safe ML pipeline and SHAP interpretability can be combined and rigorously self-audited; real patient-level data and external validation are required before any clinical inference is drawn.

David Chepkonga, A. Langat, Ebenezer Esenogho et al. · 0 citations
Conference Jul 2026

Machine Learning based Analysis and Prediction of Asthma Disease

Asthma is considered to be one of the most common chronic pulmonary diseases across the world, which calls for appropriate identification and risk assessment to enhance patients' prognosis. Machine learning approaches have proved to be effective means of analyzing patients' data to predict a disease. This paper is aimed at comparing the effectiveness of four supervised machine learning approaches, such as logistic regression, decision tree, random forest, and K-Nearest Neighbor (KNN) for predicting and analyzing the disease of asthma. The experiment involves the analysis of a big Kaggle database with 316,800 observations that include 19 clinical and demographic characteristics of patients suffering from asthma. Data was processed through cleaning and selecting particular features. Then, a 80:20 train and test splitting technique was implemented. The model effectiveness was measured with such parameters as accuracy, precision, recall, F1-score, AUC, sensitivity and specificity. As a result of experiments, KNN classifier provided the highest prediction accuracy and F1-score – 75% and 0.50 correspondingly, which proves its high ability to predict positive cases of asthma. The comparative analysis demonstrates the capabilities and drawbacks associated with each model and shows how essential it is to use suitable machine learning models in order to predict asthma risk. The suggested approach offers a cost-effective and effective way of making decisions that could be useful for healthcare practitioners.

U. Ghodeswar, P. K.Joshi, M. Keote et al. · 0 citations

Machine Learning Strategies for Malaria Risk Prediction based on Text-based Clinical Information

This study demonstrates the effectiveness of machine learning techniques in malaria risk prediction using clinical information and proposes a malaria risk prediction model using machine learning techniques based on clinical information to potentially improve early detection and treatment of malaria, ultimately reducing the burden of the disease.

Prabhat Kumar, Pragati Sahu, Smaranika Priyadarshini et al. · 2 citations
Open access Aug 2025

Ensemble Machine Learning for Malaria Diagnosis in Resource-Limited Settings Using Clinical and Demographic Features

Background Sub-Saharan Africa continues to shoulder the heaviest burden of malaria. The 2024 WHO malaria report highlighted that Africa contributed an alarming 94% of the global cases and 95% of the deaths. In the WHO African region, progress towards elimination and management of malaria is hindered by weak health systems, and lack of traditional diagnostic methods such as microscopy and malaria rapid diagnostic tests (mRDT). The primary aim of the study is to develop a machine learning (ML) ensemble model for malaria diagnosis using clinical and demographic data, tailored for resource-limited settings. Methods A retrospective study was conducted using 637 patient records from Gutu Mission Hospital and Gweru Provincial Hospital in Zimbabwe. Clinical symptoms (fever, chills, abdominal pain, headache and diarrhea) and demographic features (age, gender, residence and travel history) were analysed. Data preprocessing included handling class imbalance using Synthetic Minority Oversampling Technique (SMOTE) and feature selection using Recursive feature elimination (RFE). Seven individual ML models including Logistic regression (LR), Random Forest (RF), Decision Trees (DT), Gradient Boosting (GB), K-Nearest Neighbor (KNN), Naive Bayes (NB) and XGBoost were trained and evaluated on the malaria dataset. The individual models were further combined to build, train and evaluate ensemble models such as Bagging, Stacking, Soft Voting and AdaBoost. Model performance was assessed using accuracy, precision, confusion matrices, recall and F1score and AUR-ROC metrics. Results Clinical symptoms (chills: p=0.001, fever: p=0.003, diarrhoea: p=0.01, abdominal pain: p<0.001) were statistically significant predictors of malaria. Of the demographic factors, only travel history (p=0.02) showed significant association with malaria. Among the seven individual ML models, GB achieved the highest predictive performance (Accuracy = 0.94), followed by RF (Accuracy = 0.94%) and XGBoost (Accuracy = 0.93%). The stacking ensemble model outperformed all individual ML models and other ensemble models (bagging, soft voting and adaBoost) achieving accuracy = 0.96, precision = 0.95, recall = 0.98, F1 score= 0.96 and AUC-ROC = 0.98. Conclusion This study demonstrates that ML particularly ensemble models can be used to significantly improve malaria diagnosis. The integration of these models into a web-based application could provide a scalable and accessible diagnostic tool for healthcare workers in resource limited settings.

Panashe Nyengera, H. Takawira, F. Mlambo · 1 citation
Open access Jul 2026

A comparative analysis of machine learning and deep learning models for early oral cancer detection using multi-risk epidemiological data from multiple countries

Purpose: This study aims to evaluate the effectiveness of machine learning, deep learning, and ensemble learning methods for the early diagnosis of oral cancer using multi-risk epidemiological and behavioural data collected from several countries. The main goal is to examine whether non-diagnostic factors can support early prediction before clear clinical signs appear. Design/Methodology/Approach: The study used a dataset containing 84,922 records, including pre-diagnosis and post-diagnosis variables. Post-diagnosis variables were used only for interpretation and analysis to avoid data leakage. Several models were tested, including Random Forest, Logistic Regression, Support Vector Machine, XGBoost, LightGBM, TabNet, MLP classifier, and a voting-based ensemble model. The models were evaluated using accuracy, recall, precision, F1-score, and ROC-AUC. Research Limitation: The main limitation is that the study did not use clinical diagnostic features, medical imaging, genetic markers, or laboratory biomarkers, which may improve prediction performance. Findings: The results showed that when only non-diagnostic epidemiological and behavioural features were used, all models achieved performance close to random classification. This means that early prediction of oral cancer using these features alone is difficult. The study also showed clear regional and economic differences among countries in oral cancer prevalence, feature importance, treatment costs, and productivity losses. Practical Implication: The findings suggest that epidemiological data should be combined with clinical and biomarker-based data to develop more accurate early diagnosis systems. Social Implication: The study highlights the need for better awareness, early screening programs, and improved access to diagnostic services, especially in developing countries. Originality/Value: This study provides a realistic evaluation of oral cancer early prediction using non-diagnostic data and emphasises the importance of avoiding data leakage in medical AI studies.

G. H. Hussein, S. Elbai, K. Elayati et al. · 1 citation
Open access Aug 2026

MACHINE LEARNING MODELS FOR EARLY DETECTION AND PUBLIC HEALTH MANAGEMENT OF TUBERCULOSIS

Background: Tuberculosis (TB) continues to be one of the most pressing health issues worldwide, and in 2024, 10.6 million new cases were reported, resulting in 1.3 million deaths. However, the early and accurate diagnosis is crucial to limit transmission but the existing means are expensive, time consuming, and impractical in resource-limited environments. The objective of the present study was to build a quick and low-cost diagnostic model based on the use of blood biomarkers routinely available in the laboratory. Tuberculosis (TB) continues to be one of the more deadly diseases, particularly in the developing world. This study is aimed at developing a prototype solution where primary signs, symptoms and risk factors of TB would be identified at an early stage and machine learning (ML) predictive algorithms would be applied to them. Methods: Data for 818 confirmed TB cases and 2,618 healthy controls were analyzed in a retrospective manner. Since the dataset is imbalanced, the ROSE technique was used to balance the training set. Seven machine learning algorithms were trained, and feature selection was done using LASSO regression and forward selection to determine the most predictive variables. SHAP analysis was used to increase interpretability of the model, and the final predictive model was implemented as an interactive Shiny web application to make it easier to be used in the clinic. Result: The Gradient Boosting Machine (GBM) model demonstrated superior performance on the test set, achieving an area under the curve (AUC) of 0.821, specificity of 85.7%, and sensitivity of 64.9%. SHAP analysis highlighted platelet-to-lymphocyte ratio (PLR), monocyte-to-lymphocyte ratio (MLR), and platelet distribution width (PDW) as the most influential predictors. Adjusting the classification threshold to 0.24 improved sensitivity to 82.6% while maintaining acceptable specificity (58.9%), underscoring the model’s potential utility as a screening tool. The accompanying Shiny application enhances accessibility and practical deployment in clinical settings. Conclusion: This study presents a robust and interpretable GBM-based diagnostic model leveraging routine hematological parameters to provide a rapid, low-cost TB screening tool suitable for resource-constrained environments. The model’s high specificity may reduce unnecessary confirmatory testing, while its core predictors offer biological insights into TB-associated inflammatory processes. The interactive web application facilitates integration into clinical workflows, supporting early detection and improved TB management.

Yahya Khan, Lisa Sullivan, Danial Khan et al. · 0 citations