Skip to content
Open access

Machine learning-based prediction models for postoperative pulmonary complications in elderly patients undergoing abdominal surgery

Jul 2026 · Frontiers in Surgery · Vol 13 · 0 citations · 29 references
Medicine

TL;DR

An interpretable gradient-boosting model may support risk-stratified perioperative assessment for elderly patients undergoing abdominal surgery and Prospective multicenter validation is required before routine clinical implementation.

Abstract

Background Postoperative pulmonary complications (PPCs) are common adverse events after abdominal surgery in older adults, but existing risk scores may have limited transportability and clinical interpretability in elderly surgical populations. Methods This retrospective cohort study included 2,456 patients aged >=65 years who underwent abdominal surgery in the development/internal cohort and 542 patients in an independent external-validation cohort. Six algorithms were compared, including logistic regression, random forest, support vector machine, neural network, XGBoost, and LightGBM. Model performance was evaluated using discrimination, calibration, decision curve analysis, and external validation. SHAP analysis was used to support interpretability. Additional revision analyses examined pulmonary-function-test missingness, minor versus major PPCs, PPC co-occurrence patterns, temporal stability, COVID-era effects, and comparator-score performance. Results PPCs occurred in 425 of 2,456 patients (17.3%) in the development/internal cohort and 105 of 542 patients (19.4%) in the external-validation cohort. XGBoost showed the best overall performance, with AUCs of 0.856 (95% CI, 0.811–0.900) in the independent test set and 0.821 (95% CI, 0.781–0.861) in the external-validation cohort. At a 20% risk threshold, the independent-test sensitivity, specificity, PPV, and NPV were 87.5%, 60.9%, 32.0%, and 95.9%, respectively. SHAP analysis identified ASA physical status, COPD, upper abdominal surgery, age, emergency surgery, albumin, and surgical duration as leading contributors. The model separated patients into low-, moderate-, and high-risk groups with observed PPC rates of 6.9%, 24.4%, and 37.6%. Sensitivity analyses supported robustness to pulmonary-function missingness and temporal variation. Conclusion An interpretable gradient-boosting model may support risk-stratified perioperative assessment for elderly patients undergoing abdominal surgery. Prospective multicenter validation is required before routine clinical implementation.

Read PDF

Similar papers

Open access Jul 2026

Development and validation of a routine blood test-based model to predict in-hospital postoperative pulmonary infection in older patients with hip fracture.

BACKGROUND Postoperative pulmonary infection (PPI) is a common and serious complication in older adults undergoing hip fracture surgery, leading to prolonged hospitalization, increased costs, and increased mortality. However, simple and reliable preoperative predictors remain limited. Therefore, this study aimed to develop and validate a hematology-based machine learning model for the early prediction of PPI in older hip fracture patients. METHODS A total of 3,944 patients aged ≥ 60 years who underwent hip fracture surgery were retrospectively enrolled from three cohorts: the discovery cohort (n = 1,745, Shanghai Xuhui Central Hospital, 2016-2020), the internal validation cohort (n = 1,306, 2021-2024), and the external validation cohort (n = 893, Shanghai Putuo People's Hospital, 2016-2024). Twenty-four preoperative hematologic variables were analyzed. Six supervised machine learning algorithms were compared via fivefold cross-validation. Model performance was evaluated by the area under the receiver operating characteristic curve (AUC), accuracy, sensitivity, specificity, F1 score, calibration, and decision curve analysis (DCA). RESULTS Patients who developed PPI were generally older and exhibited a neutrophil-dominant inflammatory profile, characterized by higher white blood cell counts, neutrophil, monocyte, platelet, and C-reactive protein levels, and lower lymphocyte, eosinophil, and basophil percentages (all p < 0.001). Among the evaluated algorithms, the extreme gradient boosting (XGBoost) model achieved the best overall performance, with AUCs of 1.00, 0.96, and 0.98 in the discovery, internal, and external cohorts, respectively. Calibration curves suggested good agreement between predicted and observed probabilities, and DCA indicated favorable clinical net benefit across threshold probabilities. CONCLUSIONS A hematology-based XGBoost model was developed to predict in-hospital PPI in older adults following hip fracture surgery. The model demonstrated good discriminative performance and interpretability in this study cohort, suggesting its potential utility as a supplementary tool for cost-effective perioperative risk stratification. However, further prospective validation in diverse populations and healthcare settings is required to confirm its generalizability and clinical applicability.

Zhen Xu, Jinyu Liu, Fei Chen et al. · 0 citations
Jul 2026

Machine Learning Prediction of Postoperative Mortality in Older Emergency Surgery Patients.

INTRODUCTION The older adult population account for 35% of emergency general surgery (EGS) admissions and have an increased risk-adjusted odds of mortality compared to their younger counterparts. We aimed to create a machine learning tool utilizing deep mixture of neural networks (DMNN) to predict postoperative mortality in the older adult requiring EGS. METHODS The American College of Surgeons National Surgical Quality Improvement Program (NSQIP) 2023 database was queried with approximately one million patient records. Patients aged ≥55 y who underwent EGS were included. A DMNN model to predict postoperative mortality from patient clinical data and biomarkers was performed. Model performance was assessed using the area under the receiver operating characteristic curve, confusion matrices, and Shapley additive explanations. RESULTS A total of 91,058 patients were included of which 49,363 (54%) were females with a mean age of 71.7 y. A total of 91,058 patients were included. In the held-out test set (n=1,000; mortality rate 5.9%), the DMNN achieved an area under the receiver operating characteristic curve of 0.910 (95% confidence interval 0.878-0.938) compared with 0.936 (95% confidence interval 0.911-0.959) for the NSQIP calculator (DeLong p=0.0039). At the 0.5 probability threshold, the DMNN demonstrated substantially higher sensitivity (0.542 versus 0.119) while maintaining high specificity (0.943 versus 0.997). This resulted in fewer missed deaths (27 versus 52 false negatives) at the cost of more false positives (54 versus 3). Multithreshold and decision-curve analyses showed that relative performance is threshold-dependent. The DMNN offered advantages in sensitivity at higher thresholds, while NSQIP showed modestly higher net benefit at several lower-to-moderate risk thresholds. CONCLUSIONS Our DMNN model consisting of clinical variables performed adjunctively with the American College of Surgeons NSQIP risk calculator in perioperative mortality risk prediction for older adults undergoing EGS. This model could be further trained and utilized for improved accuracy and prediction of adverse outcomes in older EGS patients.

Asanthi M. Ratnasekera, Scott McCloud, Phillip D. Jenkins et al. · 0 citations
Aug 2026

Machine Learning Modeling for Predicting Mortality in Pediatric Patients Undergoing Elective Noncardiac Surgery: Comparison to a Regression Model.

BACKGROUND Perioperative mortality in children is relatively rare; however, accurate preoperative risk stratification is critical, as it enables anticipatory planning to mitigate the risk of death. This study aims to use machine learning (ML) to develop and internally validate a predictive model for 30-day mortality in children undergoing noncardiac surgery and compare model performance to the regression-based Pediatric Risk Assessment (PRAm) score. METHODS A retrospective study of the National Surgical Quality Improvement Program (NSQIP)-Pediatric database from 2012 to 2022, excluding 2020, was performed. Patients <18 years undergoing multispecialty surgical procedures except cardiac surgery were included. Clinically meaningful risk factors for mortality were included in the random forest and XGBoost ML models. The primary outcome was 30-day mortality. RESULTS A total of 1,023,639 unique patient encounters were included in the final analysis; 3522 (0.34%) resulted in 30-day mortality. Most patients of the 1,023,639 were ≥12 years (307,930, 30.1%), followed by 6 to 12 years (263,241, 25.7%). The majority was inpatient (596,642, 58.3%) and underwent an elective procedure (736,163, 71.9%). The most common comorbid conditions were neurologic disease (209,384, 20.5%), gastrointestinal disease (177,888, 17.4%), and central nervous system tumor or acquired abnormality (129,711, 12.7%). A total of 3.2% (33,103) were mechanically ventilated and 0.6% (6032) supported with inotropes. ML models were developed on the 70% training set (n = 716,662) and evaluated using the 30% validation set (n = 306,977). The XGBoost model demonstrated the best performance in the validation set (area under the receiver operating characteristic curve [AUC-ROC] = 0.956, area under the precision-recall curve [AUC-PR] = 0.179). The accuracy was 99.4% and the precision was 0.247, meaning that a positive prediction was associated with a 24.7% risk of mortality. The model demonstrated good calibration (Brier score = 0.003) between observed and expected probabilities. The AUC-ROC for the XGBoost model was 0.956 and for the PRAm score 0.958. There was no substantial increase in net benefit of the XGBoost model versus the PRAm score across the range of threshold probabilities. CONCLUSIONS ML can be leveraged to develop clinical prediction tools with excellent predictive performance for rare but critical outcomes like postoperative mortality. However, the ML-based models in the current study performed similarly to the regression-based PRAm score, highlighting the need to consider whether their added complexity yields meaningful clinical benefit.

S. Staffa, Eleonore Valencia, V. Tangel et al. · 0 citations
Open access Jul 2026

Machine learning-based prediction of hospital-associated complications after tibial fracture surgery in older patients: a nationwide Japanese database study.

BACKGROUND Older adults requiring emergency surgery for acute tibial fractures are vulnerable to hospital-associated complications (HACs), but admission-time risk stratification tools are lacking. We aimed to characterize HACs and develop both an ensemble prediction model and a simplified bedside risk score. METHODS This retrospective cohort study used the Japanese Diagnosis Procedure Combination database provided by JMDC Inc. Patients aged ≥ 65 years emergently admitted for tibial fracture (ICD-10: S82, 2014-2025) who underwent surgery within 5 days, with stay > 5 days, were included. The primary outcome was composite HACs during index hospitalization. Missing data were handled with multiple imputation. A Super Learner ensemble was developed and evaluated on held-out test data, and a simplified scorecard was derived using analysis of variance (ANOVA)-based feature selection and Weight-of-Evidence transformation. RESULTS Among 53,186 admissions, 5,193 met eligibility criteria. HACs occurred in 851 patients (16.4%), most commonly delirium (7.5%) and falls/trauma (5.7%). The Super Learner achieved a test-set area under the receiver operating characteristic curve (AUC) of 0.740 (95% CI 0.707-0.772), higher than conventional linear logistic regression (0.724). The most influential predictors were Hospital Frailty Risk Score, dementia, Barthel Index, comorbidity burden, and days to surgery. The simplified 7-variable scorecard achieved a test-set AUC of 0.736 (95% CI 0.703-0.769), stratifying patients into five risk groups (HAC rate: 3.45%-38.64%). CONCLUSIONS HAC risk after emergency tibial fracture surgery in older adults was driven primarily by geriatric vulnerability rather than fracture-specific factors. A simplified admission-time score may support targeted prevention, pending external validation.

Akio Shimizu, Ryota Sakamoto, H. Nagayama et al. · 0 citations
Open access Aug 2026

Early risk stratification of postoperative pneumonia after brain tumor surgery using routine perioperative variables: development and prospective multicenter validation of an interpretable prediction model

Background Early postoperative pneumonia (POP) is a common and serious complication after brain tumor surgery, but early recognition is difficult because postoperative neurological dysfunction and respiratory symptoms are often non-specific. Existing models are mostly retrospective, not designed for neurosurgical patients, and rarely prospectively validated across centers. We aimed to develop an interpretable model for early POP risk stratification. Methods We used routine perioperative data from 1,856 patients undergoing brain tumor surgery at multiple centers in China between 2022 and 2025. Ten machine learning algorithms were compared. From 41 candidate variables, 11 predictors were selected using correlation analysis and LASSO. The final locked model was prospectively tested in one internal temporal cohort and three external cohorts. Performance was assessed by AUC, calibration, and decision curve analysis. Interpretability was evaluated using SHAP, a nomogram, and a web calculator. Results Logistic regression showed the best overall performance, with an AUC of 0.897 (95% CI, 0.842–0.952) in the internal cohort and a mean AUC of 0.876 ± 0.044 across the three external cohorts. Key predictors included chronic lung disease (CLD), diabetes mellitus (DM), body mass index (BMI), admission Karnofsky Performance Status (KPS), and preoperative albumin (Alb) and glucose (Glu). Conclusion This interpretable 11-variable model enables early POP risk stratification after brain tumor surgery and may support timely preventive intervention in neurosurgical care.

H. Mao, Fengchun Mu, Xinyu Wang et al. · 0 citations