Jul 2026· Journal of Clinical Medicine· Vol 15, pp. 5339· 0 citations· 19 references
Medicine
TL;DR
Machine learning models showed acceptable performance for predicting postoperative SSI after spinal surgery, suggesting that conventional statistical approaches may remain clinically useful in structured datasets.
Abstract
Background/Objectives: Surgical site infection (SSI) remains a clinically important complication after spinal surgery. This study developed and assessed machine learning approaches for predicting postoperative SSI using routinely collected preoperative clinical variables, with emphasis on calibration and clinical applicability. Methods: In this retrospective single-center study, four prediction models were developed in patients undergoing spinal surgery: logistic regression, random forest, gradient boosting, and XGBoost. Model training used five-fold stratified cross-validation, and performance was evaluated using a hold-out internal test set. Performance was assessed using the area under the receiver operating characteristic curve (AUC), area under the precision–recall curve (AUPRC), sensitivity, precision, F1 score, Brier score, and calibration slope. SHAP analysis was performed to evaluate model interpretability. Results: The incidence of SSI was 16.6%. In cross-validation, discrimination performance was broadly comparable across models, with logistic regression showing the highest observed AUC (0.814) and AUPRC (0.484). In the hold-out test set, the same model showed the highest AUC (AUC 0.806, 95% CI 0.757–0.852) and the highest sensitivity (0.758). Calibration performance varied across models. SHAP analysis identified C-reactive protein, hemoglobin, albumin, and white blood cell count as the most influential predictors. Perioperative variables provided only modest incremental predictive value. Conclusions: Machine learning models showed acceptable performance for predicting SSI after spinal surgery. Logistic regression demonstrated performance comparable to that of the evaluated machine learning models, suggesting that conventional statistical approaches may remain clinically useful in structured datasets. Preoperative clinical and laboratory variables were the major contributors to prediction, supporting their use for routine preoperative risk stratification.
An interpretable gradient-boosting model may support risk-stratified perioperative assessment for elderly patients undergoing abdominal surgery and Prospective multicenter validation is required before routine clinical implementation.
Qiang Zhong, Guiming Huang, Wen Zhou et al.· Frontiers in Surgery· 0 citations
Surgical site infections (SSIs) are a common complication in gastrointestinal surgery, leading to major morbidity, mortality, and economic cost. There is a paucity of prediction models available for SSIs to improve the identification of patients at risk of an SSI. This review aims to evaluate the performance, validation, and methodological quality of prediction models for SSI in gastrointestinal surgery.
A systematic review was conducted of MEDLINE, Embase, and Web of Science databases from January 1, 2015, to July 3, 2025. The primary outcome was discriminative performance (area under the receiver operating characteristic curve [AUROC]). Secondary outcomes included calibration, clinical utility assessment, and validation.
From 7,692 records, 40 studies met the inclusion criteria, describing 129 distinct prediction models (86 regression-based and 37 machine learning/artificial intelligence-based). SSI incidence varied from 0.7% to 54.8%. AUROC for regression models ranged from 0.49 to 0.997 (median 0.76), and for ML/AI models from 0.50 to 0.991 (median 0.67). 27 models (20.93%) reported any form of calibration, and only 13 models (10.07%) showed a decision curve analysis.9 studies (47.50%) performed some form of external validation either of their new score and/or of a previous score, and 7 studies (17.5%) performed no validation of their newly developed score.
Contemporary SSI prediction models for gastrointestinal surgery remain characterised by inadequate validation, poor calibration reporting, insufficient assessment of clinical utility, and limited integration into electronic health records. These significant barriers must be addressed in future model development and validation to affect clinical practice.
H. Bhatti, S. Erridge, Artemis Mantzavinou et al.· British Journal of Surgery· 0 citations
A machine learning prediction tool using pre-operative clinical features to identify cases likely to require 13 or more tissue sections in Mohs micrographic surgery can accurately predict which Mohs procedures will require 13 or more sections.
Y. Aksoy, S. Lee, G. Moreno-Bonilla· medRxiv· 0 citations
Patients with lumbar spinal stenosis are typically elderly with multiple comorbidities, necessitating accurate preoperative anesthetic risk assessment. The American Society of Anesthesiologists (ASA) classification quantifies functional reserve and disease burden, serving as a widely used tool for risk stratification. However, ASA classification is often influenced by subjective factors including physician experience and varies among clinicians with different seniority, while the assessment process remains time-consuming. This study aimed to develop an automated model for anesthetic risk stratification and evaluate its performance, with the goal of providing decision support for surgical and anesthetic management in this patient population.
Clinical data of 600 patients with lumbar spinal stenosis were collected and randomly divided into training (
n
= 480) and internal validation (
n
= 120) sets. An additional 100 patients from another tertiary hospital formed an external validation set. The model was validated and hyperparameter-tuned using k-fold cross-validation. Model performance, including overall classification and high-risk identification, was evaluated using accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa, linear weighted accuracy, positive predictive value, and negative predictive value.
In the internal validation set, the accuracy, macro-average precision, macro-average recall, macro-average F1 score, weighted Kappa coefficient and linear weighted accuracy of the model are 0.97, 0.96, 0.95, 0.96, 0.93 and 0.98 respectively, while in the external validation set, they are 0.97, 0.97, 0.94, 0.96, 0.93 and 0.99 respectively. The confusion matrix heatmap shows that the error is mainly concentrated between adjacent classes, and there is no cross-class misjudgment. In the internal validation set, the model's accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and Kappa coefficient for identifying high-risk patients were 0.98, 0.95, 0.98, 0.91, 0.99, and 0.92, respectively. In the external validation set, these values were 0.98, 0.93, 0.99, 0.93, 0.99, and 0.92, respectively.
The machine learning model developed in this study demonstrates strong capability in stratifying anesthetic risk for patients with lumbar spinal stenosis, providing valuable reference for selecting surgical and anesthetic approaches.
Jitao Yang, Yixi Wang, Qihao Chen et al.· Frontiers in Medicine· 0 citations
Molecular markers emerged as the strongest predictors, whereas conventional clinical variables showed limited value, and the value of integrating molecular-genetic biomarkers with ML for personalized risk stratification and preventive care in reconstructive surgery is highlighted.
L. Sydorchuk, R. Gumennyi, Miroslav Škoda et al.· Computation· 1 citation