Skip to content
Open access

Lung cancer risk prediction using interpretable ensemble models on lifestyle and clinical data

Sep 2026 · PLoS ONE · Vol 21, pp. e0357291 - e0357291 · 0 citations · 72 references
Medicine

TL;DR

The results suggest that stacking-based ensembles can improve risk prediction from lifestyle and clinical indicators while maintaining model transparency through SHAP and LIME explanations, highlighting the potential of interpretable ensemble learning as a decision-support tool for lung cancer risk assessment.

Abstract

Lung cancer remains one of the leading causes of cancer-related mortality worldwide, where early detection and reliable risk assessment are critical for improving outcomes. This study develops and evaluates a range of ensemble learning models for lung cancer prediction using lifestyle and clinical indicators, with an emphasis on both predictive performance and interpretability. Five base learners—logistic regression, k-nearest neighbors, naïve Bayes, support vector machine, and linear discriminant analysis—were used to construct multiple boosting and bagging models. Building on these, voting and stacking ensembles were designed by selectively combining high-performing models. All approaches were evaluated on the original dataset as well as on balanced and upsampled variants derived through synthetic augmentation. Model performance was assessed using accuracy, precision, recall, F1-score, Matthews correlation coefficient (MCC), and AUC-ROC. The results show that ensemble approaches consistently outperform individual models, with voting and stacking demonstrating superior performance over boosting and bagging methods. The stacking model achieved the strongest overall performance across all evaluated models. On the original dataset, which provides a more realistic representation of practical deployment conditions, it attained an accuracy of 93.53%. Performance further improved on the balanced and upsampled datasets on the upsampled dataset. To enhance transparency, SHAP and LIME were employed to provide global and local interpretability, respectively, enabling identification of key clinical factors and patient-specific risk drivers. The analysis highlights both alignment with known clinical patterns and dataset-driven variations, supporting informed interpretation of model outputs. The results suggest that stacking-based ensembles can improve risk prediction from lifestyle and clinical indicators while maintaining model transparency through SHAP and LIME explanations. These findings highlight the potential of interpretable ensemble learning as a decision-support tool for lung cancer risk assessment.

Read PDF

Similar papers

Review Open access Aug 2026

EARLY LUNG CANCER PREDICTION USING CTGAN AND TREE-BASED MACHINE LEARNING TECHNIQUES

One of the main causes of cancer-related fatalities globally is still lung cancer, and increasing survival rates depends on early identification. In this work, a machine learning-based method for predicting lung cancer utilising clinical and lifestyle data from surveys is presented. The Synthetic Minority Over-sampling...

Bushra Khanam, Lubna Nausheen, F. Fatima · 0 citations
Open access Sep 2026

Explainable ensemble learning framework with recursive feature elimination for lung cancer stage prediction using demographic and lifestyle parameters

The proposed explainable ensemble learning framework provides accurate, robust, and interpretable lung cancer stage prediction through the integration of leakage-free model development, optimized ensemble learning, explainable artificial intelligence, statistical validation, and component-wise ablation analysis.

M. Alfuraydan, Shahid Mohammad Ganie, Ehab Seedahmed et al. · 0 citations
Conference Open access 2026

Comparing Polynomial Logistic Regression and XGBoost for Structured Lung Cancer Risk Prediction Using Demographic, Behavioral, and Symptom-Based Variables

The Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and early risk prediction may support timely screening and intervention. This study compared an interpretable regression-based model with a more flexible machine learning model for structured lung cancer prediction. The publicly av...

Zhi-Xiang Li · 0 citations
Aug 2026

Comparison and interpretability of machine learning algorithms to predict survival of patients with prostate cancer.

OBJECTIVE To evaluate the performance and interpretability of multiple ML algorithms for survival prediction in prostate cancer (CaP) and to assess whether these approaches can improve upon established clinical risk prediction models. METHODS Using the Surveillance, Epidemiology, and End Results (SEER) 17 Database, w...

Isaac E. Kim, Cecile P. G. Meier-Scherling, S. Zhou et al. · 0 citations

Cancer Risk Prediction Based on Integrated Machine Learning with Optimal Threshold Selection

The study suggests that this framework can serve as a clinical screening tool, prioritizing imaging examinations for high-risk individuals and using liquid biopsy for re-screening of moderate-risk groups, and driving the translation of precision cancer prevention from theory to practice.

Wei-Pei Liu · 0 citations
#explainable ai Open access Sep 2026

An XGBoost–logistic regression hybrid stacking model for breast cancer outcome prediction

Accurate prediction of breast cancer outcomes remains challenging due to high-dimensional, imbalanced clinical datasets and the need to balance predictive performance, calibration, interpretability, and model simplicity. We propose an AICc-guided hybrid stacking framework that integrates Logistic Regression...

Oyetayo Oyebisi, W. Chacha, F. Amoyedo · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.