A Domain-Structured Ensemble Framework for Perioperative Outcome Prediction Using Electronic Health Record Data
Abstract
Perioperative risk prediction models are often limited by narrow surgical populations, incomplete intraoperative data, poor calibration, and limited interpretability. We present a domain-structured ensemble framework for perioperative outcome prediction using routinely collected electronic health record (EHR) data, designed for extensibility across diverse clinical endpoints. The framework organizes predictors into three clinically motivated domains: patient-related (baseline vulnerability), surgery-related (procedural characteristics), and anesthetics-related (intraoperative exposures and physiologic perturbations). Domain-specific gradient boosting models generate independent risk estimates, which are integrated through a logistic regression meta-learner to produce calibrated predictions. We demonstrate the framework using postoperative delirium (POD) as an exemplar application in a case–control sample of 5,386 surgical encounters (2,693 cases, 2,693 controls) from a statewide health information exchange. POD was identified through dual confirmation requiring both ICD codes and positive Confusion Assessment Method screening within seven postoperative days; patients with preexisting dementia were excluded. The stacked meta-learner achieved an area under the receiver operating characteristic curve (AUROC) of 0.899 (95% CI: 0.891–0.906), precision-recall AUC of 0.881, and Brier score of 0.126, compared with 0.849 for the best single-stage model. Domain ablation analysis confirmed that the three-domain architecture improved both discrimination and calibration relative to a surgery-only model (AUROC 0.879, Brier 0.140). Temporal validation on held-out post-2017 data yielded AUROC of 0.915. Calibration was excellent: intercept −0.006 (95% CI: −0 083 to 0.070), slope 1.035 (95% CI: 0.982 to 1.088). Decision curve analysis, corrected for case–control sampling, demonstrated positive net benefit across clinically plausible risk thresholds. The framework’s modular architecture supports substitution of outcome definitions, extension of predictor domains, and dynamic risk updating as perioperative data accrue. With external validation, this approach may serve as a scalable foundation for interpretable, calibration-aware perioperative clinical decision support across multiple postoperative outcomes.