Skip to content

Operative-Time Gradients Reveal Miscalibration in Microsurgical Risk Prediction.

Aug 2026 · Journal of reconstructive microsurgery · 0 citations
Medicine

TL;DR

Generalized ACS NSQIP morbidity prediction was accurate in aggregate but systematically miscalibrated across the operative-time gradient, overestimating risk in shorter operations and underestimating it in longer operations.

Abstract

Background

General surgical risk models support perioperative counselling, but acceptable overall calibration may conceal clinically important error. We evaluated generalized American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) predictions after microsurgical reconstruction and tested whether morbidity calibration varied across operative time.

Methods

Adult ACS NSQIP cases from 2014-2023 were analyzed. Plastic-surgery-service cases meeting exact principal Current Procedural Terminology (CPT) or principal-procedure-text criteria formed the primary cohort. Performance was assessed using observed-to-expected (O:E) ratios, discrimination, and calibration measures. Operative-time calibration was examined by quartiles, deciles, and continuous spline models. Sensitivity analyses used exact principal CPT codes alone and excluded cases with a known morbidity component or reoperation on postoperative day 0 or 1.

Results

Among 20,604 microsurgery cases, observed and predicted morbidity were similar overall (12.52% vs. 12.79%; O:E, 0.98; 95% confidence interval [CI], 0.94-1.02), although discrimination was modest (area under the receiver operating characteristic curve, 0.639; 95% CI, 0.627-0.651). Mortality was uncommon (21 deaths; O:E, 0.81; 95% CI, 0.50-1.24). Aggregate calibration concealed a graded reversal across operative time. Morbidity was overpredicted in the shortest quartile (8.66% observed vs. 12.87% predicted; O:E, 0.67; 95% CI, 0.61-0.74) and underpredicted in the longest quartile (16.94% vs. 12.96%; O:E, 1.31; 95% CI, 1.22-1.40). The prespecified linear interaction was nonsignificant (P = 0.183), whereas flexible continuous analyses demonstrated calibration variation before and after global recalibration (both P < 0.001). Findings persisted in both sensitivity analyses.

Conclusions

Generalized ACS NSQIP morbidity prediction was accurate in aggregate but systematically miscalibrated across the operative-time gradient, overestimating risk in shorter operations and underestimating it in longer operations. Realized operative time should be interpreted as a postoperative marker of incompletely captured procedural complexity and intraoperative course, not as a causal exposure or preoperative predictor.

View source

Similar papers

Aug 2026

Thirty-day outcomes after bariatric conversion operations: a national analysis of 94,828 cases.

BACKGROUND Conversion bariatric operations are increasing and now comprise nearly 10% of procedures nationally, yet their 30-day risk compared with primary operations remains poorly characterized at the pathway level. OBJECTIVES To estimate the adjusted excess 30-day serious adverse event (SAE) risk of conversion versus primary bariatric operations of the same final anatomy, and to identify which pathways carry the highest risk. SETTING National Metabolic and Bariatric Surgery Accreditation and Quality Improvement Program (MBSAQIP) Participant Use File (PUF), 2020 to 2024. METHODS We performed anatomy-stratified logistic regression with marginal standardization on 949,507 bariatric operations (94,828 conversions). The primary outcome was a 17-component composite SAE. Adjusted risk differences (RDs) per 1000 were computed for sleeve gastrectomy (SG), Roux-en-Y gastric bypass (RYGB), and biliopancreatic diversion with duodenal switch (DS)/single-anastomosis duodeno-ileal (SADI) strata, and a within-conversion model identified associated factors. RESULTS Conversion was associated with higher adjusted 30-day SAE risk across all anatomies: SG (RD 18.2 per 1000 [95% confidence interval, CI: 15.0-21.3], number needed to harm [NNH] 55), RYGB (22.1 [19.5-24.7], NNH 45), and DS/SADI (17.5 [9.4-25.7], NNH 57). The excess was concentrated in utilization events (reoperation, reintervention, and readmission) and selected technical complications, such as leak, organ-space infection, and bowel obstruction. Among conversions, a prior RYGB carried the highest risk (adjusted odds ratio [aOR]: 2.54 versus a prior sleeve). The excess showed no detectable change across study years. CONCLUSIONS Conversion operations carry a clinically meaningful excess 30-day SAE risk that is greatest after a prior gastric bypass. Conversion status should be treated as a distinct risk category rather than benchmarked against primary procedures alone.

R. Dallal, Priscilla Lam, Aditya Das · 0 citations
Aug 2026

American College of Surgeons NSQIP Hospital Benchmarking Using Bayesian Variational Inference to Adjust for Many CPT Codes.

BACKGROUND Current ACS NSQIP benchmarking relies primarily on the principal Current Procedural Terminology (CPT) code for procedure-related risk adjustment because conventional regression methods cannot efficiently accommodate numerous sparse CPT variables. Although CatBoost (CATB) can incorporate multiple CPT codes to improve risk estimation, it is computationally intensive and may produce unstable estimates for infrequently performed procedures. This study evaluated Bayesian automatic differentiation variational inference (ADVI) as an alternative approach. STUDY DESIGN CATB and ADVI were applied to 2019-2023 ACS NSQIP data to estimate CPT-based risk adjustment using all available CPT codes (up to 21 per patient) across 31 outcomes. Performance was compared using calibration (Hosmer-Lemeshow statistic), discrimination (area under the receiver operating characteristic curve), computational time, and effects on hospital-level morbidity benchmarking relative to principal CPT-only risk adjustment. RESULTS Both approaches demonstrated similar discrimination across all 31 outcomes, whereas ADVI consistently achieved superior calibration and markedly lower computational time, with the greatest advantages observed in larger datasets. Among the 15 largest models, mean Hosmer-Lemeshow statistics were 125.23 for ADVI versus 764.72 for CATB, while mean computation times were 9.8 versus 1447.7 minutes, respectively. ADVI and CATB produced similar, although not identical, hospital benchmarking results. CONCLUSIONS ADVI is a computationally efficient and statistically robust alternative to CATB for incorporating multiple CPT codes into ACS NSQIP risk adjustment. Improved calibration and more stable estimation for sparse procedure combinations may enhance the reliability and scalability of procedure-based benchmarking.

Yaoming Liu, Mark E. Cohen, A. Grieco et al. · 0 citations
Aug 2026

Surgeon Characteristics Associated With Average CPT® 31237 Utilization per Medicare Beneficiary.

BackgroundA retrospective cohort study of 69,170 patients who underwent endoscopic sinus surgery between 2015 and 2022 found that 38.7% of the variation in the number of postoperative debridements was explained at the surgeon level. The extent to which postoperative debridements patterns after endoscopic sinus surgery depended on the surgeon motivates further research about relevant surgeon characteristics.ObjectiveTo identify surgeon characteristics associated with average Current Procedural Terminology (CPT®) 31237 utilization per Medicare beneficiary.MethodsWe performed a retrospective cross-sectional study using summary Medicare data from 2023, filtering for CPT 31237 ("[b]iopsy or removal of nasal polyp or tissue using an endoscope"). We hypothesized that where surgeons train, whether they complete a fellowship in rhinology, or how much they are reimbursed is associated with average CPT 31237 utilization per beneficiary. We performed multivariable linear regression to analyze surgeon characteristics.ResultsThe sample included 1016 surgeons. Their sex distribution was 86.8% male and 13.2% female, and their mean years of postgraduate experience was 22.6 (SD 10.3). Of the 33.4% of surgeons who completed a fellowship, the majority trained in rhinology (71.4%), followed by facial plastics (14.8%). Approximately 96.5% of surgeons had average CPT 31237 utilization per beneficiary of 3 or less. The median number of beneficiaries per surgeon was 18 (IQR 13-28). Residency or fellowship program did not explain variation in average CPT 31237 utilization per beneficiary. Completing fellowship in rhinology was associated with 0.13 (95% CI 0.01-0.25, P = .027) greater average CPT 31237 utilization per beneficiary. Each additional $100 in the average standardized amount paid for CPT 31237 by Medicare was associated with 0.32 (95% CI 0.20-0.43, P < .001) greater average CPT 31237 utilization per beneficiary.ConclusionAverage CPT 31237 utilization per beneficiary may be associated with the type of fellowship which surgeons complete and the amount which they are reimbursed.

Max J. Hyman, Christopher M. Low, Jayant M. Pinto et al. · 0 citations
Review Aug 2026

Impact of Structural Anomalies and Anatomical Variations on Operative Time in Abdominal Surgery-Systematic Review.

Prolonged operative time is associated with increased surgical complications, resource utilization, and cost. While factors such as surgeon experience and patient habitus are well recognized, the impact of anatomical anomalies and variations remains poorly quantified. This systematic review evaluates the influence of abdominal structural anomalies and anatomical variations on operative time. Following PRISMA 2020 guidelines, a search of PubMed (2000-2025) identified 82 eligible studies. Operative-time impact was assessed using added operative time relative to reference benchmarks. The unit of analysis was the procedure-level observation. Observations were stratified by procedure category and surgical complexity and summarized using median added operative time and range. Among procedure-level observations with available benchmarks, 13 of 14 (92.9%) structural anomalies, 48 of 54 (88.9%) situs inversus cases, and 10 of 11 (90.9%) other anatomical variants demonstrated increased operative time relative to baseline. Structural anomalies were associated with larger median increases, particularly in colorectal procedures, whereas anatomical variations demonstrated a more heterogeneous effect across procedural systems. Higher-complexity procedures, including pancreatic operations, showed greater absolute increases in operative time. Preoperative identification of anatomical differences was associated with attenuation of operative delays. Structural anomalies are frequently associated with increased operative duration, whereas anatomical variations demonstrate a more variable and context-dependent effect. The magnitude of operative time impact is influenced by procedural complexity and preoperative planning. Incorporating anatomical variability into surgical planning may reduce avoidable operative delays and improve operative efficiency.

C. Alleyne, M. Montalbano, Kazzara Raeburn et al. · 0 citations
Open access Aug 2026

AI-based disease severity grading predicts complications in laparoscopic appendectomy.

BACKGROUND Appendicitis severity underpins contemporary management guidelines, where laparoscopic appendectomy remains gold standard. Preoperative measures poorly predict actual disease severity or complication risk, while operative grading systems such as the American Association for the Surgery of Trauma (AAST) remains largely confined to research settings. Artificial intelligence (AI) may provide practical solutions. We evaluated a previously validated AI-derived surgical video assessment of disease severity for predicting perioperative complications. METHODS This retrospective study included consecutive surgical videos (6/2022-1/2024) routinely analyzed by the AI platform. AI-derived severity scores were stratified into Low (uncomplicated) and High (complicated) groups. Multivariable analysis identified independent predictors of complications. Operative-AAST served for benchmarking. Model discrimination (AUC), post hoc recalibration plot, and decision curve analysis (DCA) were evaluated. RESULTS Of 632 cases, 74.5% were low severity and 25.5% high. The High group had higher complication rates (26.7% vs. 10%; p<0.001), including intraoperative (11.8 vs. 3.2%; p < 0.001) and postoperative complications (15.5 vs. 7.5%; p = 0.005). AI-derived severity independently predicted complications (OR 2.76, 95% CI 1.62-4.73; p < 0.001), even after adjustment for operative-AAST (OR 1.90, 95% CI 1.07-3.36; p = 0.028). Discrimination was modest (AUC = 0.63), similar to operative-AAST (AUC = 0.68). Calibration plot showed incremental increase in complication rates across probabilities, with acceptable agreements at extremes and improved alignment at intermediate-risk. DCA showed the model had highest net benefit at intermediate-risk, comparable or higher than operative-AAST. CONCLUSIONS Automated AI-based surgical video assessment shows promise as a complementary tool for risk-prediction of laparoscopic appendectomy. It offers scalable risk stratification that may be implemented in routine clinical practice. Nevertheless, further study and model refinement are warranted.

Tal Kardish, M. Ortenzi, E. Nizri et al. · 0 citations
Open access Aug 2026

P1.064. Limitations of Preoperative Prediction Models for Complications After Esophagectomy: A Multi-Center Analysis

Esophageal Cancer: Surgical Treatment of Esophageal Cancer Preoperative risk stratification for esophagectomy complications relies on clinical prediction models; however, their discriminative performance in multi-institutional settings remains poorly defined. We hypothesized that standard preoperative variables would demonstrate limited predictive validity across heterogeneous surgical cohorts. We analyzed 2,490 patients undergoing esophagectomy across four institutions in Asia (total n=2,490; individual center n range 75–1,012). Four outcomes were studied: recurrent laryngeal nerve palsy (RLNP), anastomotic leak (AL), pulmonary complications (PC), and vocal cord palsy (VCP). Logistic regression models with bootstrap-validated odds ratios (1,000 iterations) were evaluated by 5-fold cross-validated AUC. SHAP (SHapley Additive exPlanations) via Gradient Boosting Machines quantified variable importance. Decision curve analysis (DCA) assessed net clinical benefit across threshold probabilities 2–70%. Association between tumor location and each complication was assessed using chi-squared tests. All prediction models demonstrated poor-to-fair discrimination: RLNP AUC 0.533, AL AUC 0.586, PC AUC 0.676, and VCP AUC 0.556. Tumor location was the only statistically significant categorical predictor of RLNP—upper/cervical location was associated with higher RLNP incidence compared to middle thoracic tumors (39.1% vs. 28.0%; OR 1.22, 95%CI 1.05–1.42; p=0.009). No significant association was observed between tumor location and AL, PC, or VCP. DCA demonstrated negligible clinical net benefit for RLNP and AL models; only the PC model provided modest benefit (max net benefit gain +0.057) at threshold probabilities of 5–20%. SHAP analysis identified FEV1%, PNI score, and BMI as the highest-importance variables for RLNP prediction, with tumor location ranking sixth—indicating that location contributes a statistically real but clinically modest signal. Standard preoperative variables are insufficient for individualized risk stratification of RLNP, anastomotic leak, or vocal cord palsy after esophagectomy. Statistical significance (p=0.009 for location–RLNP association) does not translate to clinically meaningful predictive power (AUC 0.533). Tumor location should be incorporated into RLNP preoperative counseling. Improved prediction will require prospective integration of real-time intraoperative data. Pulmonary complication risk approaches clinically actionable prediction (AUC 0.676) and may guide respiratory prehabilitation targeting.

Si-miao Lu, Yi Zhu, Yong-tao Han et al. · 0 citations