MOY 01 Prediction Models for Surgical Site Infections in Gastrointestinal Surgery: A Systematic Review of Regression and Artificial Intelligence Approaches
Aug 2026· British Journal of Surgery· Vol 113· 0 citations
Abstract
Surgical site infections (SSIs) are a common complication in gastrointestinal surgery, leading to major morbidity, mortality, and economic cost. There is a paucity of prediction models available for SSIs to improve the identification of patients at risk of an SSI. This review aims to evaluate the performance, validation, and methodological quality of prediction models for SSI in gastrointestinal surgery.
A systematic review was conducted of MEDLINE, Embase, and Web of Science databases from January 1, 2015, to July 3, 2025. The primary outcome was discriminative performance (area under the receiver operating characteristic curve [AUROC]). Secondary outcomes included calibration, clinical utility assessment, and validation.
From 7,692 records, 40 studies met the inclusion criteria, describing 129 distinct prediction models (86 regression-based and 37 machine learning/artificial intelligence-based). SSI incidence varied from 0.7% to 54.8%. AUROC for regression models ranged from 0.49 to 0.997 (median 0.76), and for ML/AI models from 0.50 to 0.991 (median 0.67). 27 models (20.93%) reported any form of calibration, and only 13 models (10.07%) showed a decision curve analysis.9 studies (47.50%) performed some form of external validation either of their new score and/or of a previous score, and 7 studies (17.5%) performed no validation of their newly developed score.
Contemporary SSI prediction models for gastrointestinal surgery remain characterised by inadequate validation, poor calibration reporting, insufficient assessment of clinical utility, and limited integration into electronic health records. These significant barriers must be addressed in future model development and validation to affect clinical practice.
BACKGROUND
Pre-operative assessment for postoperative infection risk helps identify patients for personalised decision-making and management.
OBJECTIVE
This systematic review evaluates existing prediction models for infection, focusing on their validation and implementation status.
DATA SOURCES AND ELIGIBILITY CRITERIA
PubMed, Embase and the Cochrane Library were searched for studies on the development, validation- and implementation of multivariable models utilising pre-operative predictors to estimate the risk of postoperative infections within 30-days of elective, major noncardiac, non-intracranial surgery.
RESULTS
Of 151 included studies, 88 reported model development (267 distinct models), 88 assessed model validity (314 validation analyses), and none described implementation. Models predominantly predicted surgical site infections (SSI, n = 88), pneumonia (n = 45) and general (unspecified) infections (n = 57). Age (66%), sex (53%) and ASA score (48%) were the most common predictors. The American College of Surgeons Surgical Risk Calculator (ACS SRC) and SUrgical Risk Pre-operative Assessment System (SURPAS) were most frequently validated, with 225 and 35 external validations respectively. Reported c-statistics of the ACS SRC were median 0.61 [range 0.43 to 0.85], 0.66 [0.44 to 0.95] and 0.64 [0.31 to 0.97] for prediction of SSI, pneumonia, and urinary tract infection (UTI), respectively. For SURPAS, median c-statistics were 0.62 [range 0.52 to 0.78] and 0.59 [0.52 to 0.82] for UTI and general infection. Overall, most studies scored high risk of bias.
CONCLUSIONS
Out of 267 prediction models for postoperative infections identified, ACS SRC and SURPAS were most frequently validated. However, the clinical utility of even these models is limited because of poor and highly variable predictive performance and low methodological quality of validation studies.
N. de Mul, Katja M Scheffer-Wesdorp, D. Verlaan et al.· European Journal of Anaesthe...· 0 citations
Background To systematically evaluate the quality and performance of predictive models for postoperative infection risk following hip fractures, to identify reliable tools for clinical practice and provide an evidence-based foundation for the development of higher-quality predictive models in the future. Methods A systematic search was conducted on nine databases to retrieve relevant publications, from their inception up to 1 February 2026. Two researchers independently screened the literature and extracted data. They assessed the model bias and applicability using the Predictive Model Risk of Bias Assessment Tool (PROBAST) and the Checklist for Reporting on Multivariate Predictive Models for Individual Prognosis or Diagnosis-Artificial Intelligence (TRIPOD+AI). Results A total of 17 articles were included, covering 21 predictive models, with postoperative infection rates ranging from 1.61 to 24.56%. A meta-analysis of 11 high-frequency predictive factors revealed that hypoproteinemia, diabetes, pulmonary disease, ASA classification, smoking, indwelling catheter duration, age, and surgical duration were independent risk factors, while gender and albumin were not statistically significant. Furthermore, the area under the curve (AUC) for the included models ranged from 0.699 to 0.946. While most models performed well, all 17 studies were rated as having a high risk of bias by PROBAST, and the reporting quality of all studies according to TRIPOD+AI was relatively low, primarily due to retrospective study designs, regional bias, inadequate data analysis, insufficient external validation, and a lack of transparency in the research process. Conclusion Current predictive models generally demonstrate good overall predictive performance; however, most models suffer from issues such as single-center development, insufficient external validation, and methodological limitations. In the future, more multicenter, large-sample prospective studies should be conducted, and strategies for variable handling and model validation should be optimized to improve the generalizability and clinical translation of predictive models. Systematic review registration https://www.crd.york.ac.uk/PROSPERO/view/CRD420261289616, identifier (CRD420261289616).
Jiye Pan, Juan Shi, Ya-ting Ai et al.· Frontiers in Medicine· 0 citations
Esophageal Cancer: Other
Esophagectomy is associated with substantial morbidity, mortality, and resource utilization. Traditional regression-based risk tools may inadequately capture complex nonlinear interactions. Contemporary evidence on machine learning (ML) models predicting postoperative outcomes after esophagectomy was synthesized, focusing on discrimination, validation, and comparison with conventional regression approaches.
A PRISMA-guided systematic review was conducted using Embase, MEDLINE, PubMed, and the Cochrane Library in January 2026. A total of 196 studies were identified. After title and abstract screening, 32 studies underwent full-text review, of which 10 met final inclusion criteria as ML-focused prediction models in esophagectomy populations. Extracted data included study design, cohort size, procedure type, predicted outcome, modeling approach (ML versus regression), validation strategy (internal or external), performance metrics (e.g., area under the receiver operating characteristic curve [AUROC]), and reporting elements such as calibration and decision-curve analysis. ML-focused studies were defined as those applying algorithms including gradient boosting, support vector machines, neural networks, or survival forests to postoperative outcome prediction.
Ten studies applying ML models to esophagectomy outcomes were included (median cohort size 700; range 200–4700). Anastomotic leak was the most frequently predicted outcome (4/10), followed by mortality, major complications, readmission, strictures, and recurrence or survival. Common algorithms included gradient boosting (XGBoost, LightGBM, GBM), support vector machines, neural networks, and survival forests. Reported discrimination ranged from moderate to high (AUROC 0.64 for 90-day mortality and 0.65–0.70 for major complications, increasing to 0.79–0.90 for anastomotic leak prediction; some internally validated models reported AUROC >0.95). Three studies performed independent external validation, and performance generally declined in external cohorts. Comparative analyses demonstrated that ML often matched but did not consistently outperform regression-based models.
Machine learning models for postoperative risk prediction after esophagectomy demonstrate promising discrimination, particularly for anastomotic leak. Although external validation remains limited, ML approaches are still in early stages of clinical translation. With prospective data integration and robust multicenter validation, ML has the potential to enhance individualized risk stratification, support shared decision-making, guide perioperative planning, and improve allocation of postoperative resources in esophageal surgery.
T. Wang, Otari Beldishevski-Shotadze, N. Evennett· Diseases of the esophagus· 0 citations
Artificial intelligence (AI) is increasingly applied in clinical practice to enhance prediction of postoperative outcomes. This systematic review evaluated the performance and clinical relevance of AI-based prognostic models for patients undergoing percutaneous nephrolithotomy (PCNL). A comprehensive search of PubMed, Embase, Scopus, Web of Science, and Google Scholar was conducted on 7 August 2025 to identify original studies that used AI to predict outcomes such as stone-free status. Risk of bias was assessed using the Prediction model Risk of Bias Assessment Tool-Artificial Intelligence extension. A narrative synthesis of study characteristics, and reported performance metrics of AI models was conducted. Twenty-one studies involving 24,087 patients met the inclusion criteria. Tree-based, support vector, discriminant, and regression models achieved the highest median AUCs (0.79-0.81) for predicting stone-free status, whereas neural networks showed lower performance (median 0.60). Tree-based and similarity-based models performed best for predicting the need for adjuvant therapy. Neural networks performed well for bleeding-related outcomes (median AUC 0.87), while tree-based and regression models showed more consistent performance for procedural complications and hospitalization. Infection-related outcomes were predicted most accurately by tree-based and neural network models (median AUCs 0.89-0.90). Temporal trends revealed a shift from early reliance on neural networks and support vector machines to increased use of tree-based approaches in recent years. Overall, AI models offer useful support for perioperative decision-making, though their reliability varies across outcomes. Model performance is highest for stone-free status and infection-related outcomes, but lower for rare or poorly-defined endpoints such as bleeding or peri-procedural complications, reflecting limitations in available features and outcome heterogeneity. Clinically, current models may inform management but are not yet sufficient to dictate care. Future research should focus on standardized outcome definitions, multi-institutional datasets, and integration of high-dimensional data such as imaging and radiomics to develop robust, externally validated predictive tools for prospective implementation in PCNL practice.
Objective: Cardiac surgery carries a significant risk of complications and mortality. Artificial intelligence (AI), particularly machine learning (ML), is increasingly being explored to enhance perioperative risk prediction and support clinical decision-making. This systematic review evaluates the clinical applications, predictive performance, and limitations of AI models in cardiac surgery. Methods: PubMed and Embase were searched for studies published between January 2020 and July 2025. Of 939 records identified, 178 studies met the inclusion criteria following screening and full-text review. Included studies applied AI to predict clinical outcomes in patients undergoing cardiac surgery. Key outcomes assessed were model performance metrics and their clinical utility. Results: Among the 178 included studies, 114 (64%) were conducted in the United States or China. Most studies (n = 168, 94%) used retrospective designs and focused on adult populations. Random forest (n = 82, 46%), logistic regression (n = 82, 46%), and eXtreme Gradient Boosting (n = 70, 39%) were the most frequently used algorithms. AI applications primarily targeted the prediction of postoperative complications (n = 102, 57%) and mortality (n = 70, 39%), with common outcomes including acute kidney injury and stroke. ML models consistently outperformed traditional clinical risk scores (n = 39). SHapley Additive exPlanations was the most common interpretability method (n = 66, 37%). Only 26% of studies included external validation, and just 19% adhered to TRIPOD guidelines. Conclusions: AI models demonstrate superior predictive performance in cardiac surgery compared with traditional risk scores, but concerns regarding validation, transparency, and generalizability must be addressed to enable implementation. Graphical Abstract This is a visual representation of the abstract.
Background and objective Intestinal obstruction is a common surgical emergency in which delayed recognition of strangulation, ischemia, or failure of non-operative management can lead to bowel necrosis, sepsis, and death. Prediction models have been developed for several related but distinct tasks, including diagnosis of obstruction, prediction of urgent surgery, prediction of strangulation or irreversible ischemia, and prediction of failure of conservative treatment. This review critically summarizes these models and clarifies their clinical scope, validation status, and translational limitations. Methods PubMed was searched for studies published from database inception to October 31, 2025, with the initial search performed during manuscript preparation on December 1, 2025 and the final update search conducted on June 18, 2026. Search terms were related to intestinal obstruction, small bowel obstruction, prediction model, nomogram, scoring system, artificial intelligence, machine learning, deep learning, and computed tomography. Eligible articles reported or discussed predictive, diagnostic, or decision-support models for intestinal obstruction. We excluded papers without model-related content, non-clinical mechanistic studies unless used to explain modeling rationale, case reports, and articles not providing sufficient methodological or performance information. Because the evidence was heterogeneous in population, endpoint, modality, and design, the review was synthesized narratively rather than meta-analyzed. Key findings Conventional clinical scores remain attractive because they are transparent, inexpensive, and quickly calculable, but their performance varies across endpoints and settings. CT-integrated models improve anatomical and ischemic risk assessment but depend on imaging availability and reader expertise. Machine-learning and deep-learning models, including multimodal systems combining electronic health records and imaging, have reported high discrimination in selected datasets; however, many studies remain retrospective, single-center, and incompletely externally validated. Therefore, AUC values across studies should not be interpreted as directly comparable evidence of superiority. Conclusions The field is moving from static, single-modality scores toward dynamic, multimodal decision-support systems. The most immediate research priorities are prospective multicenter validation, standardized endpoint definitions, calibration and decision-curve reporting, explainability, fairness assessment, privacy-preserving data sharing, and workflow integration in emergency surgical care.
Zehao Liu, Li He, Qiangqiang Zhang et al.· Frontiers in Surgery· 0 citations