Aug 2026· Frontiers in Endocrinology· Vol 17· 0 citations· 28 references
Medicine
TL;DR
Findings suggest that prediction of amputation level is feasible, validated in a temporally separated cohort, and clinically interpretable, and may support future decision-support applications, although further validation is required before clinical implementation.
Abstract
Background Accurate preoperative prediction of whether an initially limb-preserving strategy in diabetic foot management will culminate in minor or major amputation remains a clinical challenge. This study aimed to develop and evaluate using two temporally separated cohorts a machine-learning framework using routinely available baseline clinical, laboratory, and selected imaging and vascular variables. Methods Two temporally separated cohorts were used, with Dataset 1 for model development and Dataset 2 for temporally separated evaluation. A 20-repetition stratified outer-split workflow was implemented, incorporating two-step feature selection, Optuna-based hyperparameter optimization, training-only SMOTE, and threshold tuning to maximize the F2-score under a recall constraint of ≥0.70. Six classifiers were evaluated using average precision (AP), ROC-AUC, recall, precision, specificity, accuracy, and Brier score. Results The major-amputation group exhibited a more severe baseline phenotype, including higher inflammatory burden, worse neuropathy and wound severity, and a higher prevalence of necrotizing fasciitis. Internally, multilayer perceptron achieved the highest AP (55.9% ± 13.6%). In external evaluation, k-nearest neighbors achieved the highest AP (65.1% ± 10.2%) and recall (72.8% ± 19.6%), whereas multilayer perceptron showed higher precision and specificity. Key contributors included necrotizing fasciitis, neuropathy severity, hemoglobin, PEDIS classification, and inflammatory indices. Conclusion These findings suggest that prediction of amputation level is feasible, validated in a temporally separated cohort, and clinically interpretable, and may support future decision-support applications, although further validation is required before clinical implementation.
BACKGROUND
Older adults requiring emergency surgery for acute tibial fractures are vulnerable to hospital-associated complications (HACs), but admission-time risk stratification tools are lacking. We aimed to characterize HACs and develop both an ensemble prediction model and a simplified bedside risk score.
METHODS
This retrospective cohort study used the Japanese Diagnosis Procedure Combination database provided by JMDC Inc. Patients aged ≥ 65 years emergently admitted for tibial fracture (ICD-10: S82, 2014-2025) who underwent surgery within 5 days, with stay > 5 days, were included. The primary outcome was composite HACs during index hospitalization. Missing data were handled with multiple imputation. A Super Learner ensemble was developed and evaluated on held-out test data, and a simplified scorecard was derived using analysis of variance (ANOVA)-based feature selection and Weight-of-Evidence transformation.
RESULTS
Among 53,186 admissions, 5,193 met eligibility criteria. HACs occurred in 851 patients (16.4%), most commonly delirium (7.5%) and falls/trauma (5.7%). The Super Learner achieved a test-set area under the receiver operating characteristic curve (AUC) of 0.740 (95% CI 0.707-0.772), higher than conventional linear logistic regression (0.724). The most influential predictors were Hospital Frailty Risk Score, dementia, Barthel Index, comorbidity burden, and days to surgery. The simplified 7-variable scorecard achieved a test-set AUC of 0.736 (95% CI 0.703-0.769), stratifying patients into five risk groups (HAC rate: 3.45%-38.64%).
CONCLUSIONS
HAC risk after emergency tibial fracture surgery in older adults was driven primarily by geriatric vulnerability rather than fracture-specific factors. A simplified admission-time score may support targeted prevention, pending external validation.
Akio Shimizu, Ryota Sakamoto, H. Nagayama et al.· Aging Clinical and Experimen...· 0 citations
Background/Objectives: This study aimed to develop machine learning models—specifically logistic regression (LR), random forest (RF), and deep neural network (DNN) models—using initial and 1-month post-stroke clinical data to predict 6-month upper and lower extremity motor functional outcomes in patients with ischemic stroke. Additionally, we sought to evaluate the potential improvement in discriminative performance and clinical utility achieved by integrating 1-month reassessment data. Methods: We analyzed retrospective cohort data from 353 patients with ischemic stroke. Two prediction models were constructed: (1) Model 1, which used only early-stage clinical data, and (2) Model 2, which incorporated both early-stage and 1-month post-stroke clinical data. Model performance and clinical utility were evaluated using the area under the receiver operating characteristic curve (ROC-AUC), DeLong’s test, calibration analysis, decision curve analysis (DCA), and variable importance analysis. Results: Although Model 2, which incorporated 1-month data, generally showed an upward trend in discriminative performance across all models for both upper and lower extremity prediction compared to Model 1, a statistically significant improvement was observed only in the LR model for upper extremity prediction (test AUC increased from 0.889 to 0.990; ΔAUC = +0.102, p = 0.037). For all other models—including the RF and DNN models for the upper extremity, as well as all lower extremity prediction models—the observed increases in AUC did not reach statistical significance according to DeLong’s test. In calibration analyses, the LR model exhibited the most stable calibration for both extremities. In DCA, Model 2 generally yielded a higher net benefit across most threshold probability ranges compared to Model 1 than Model 1 across most threshold probability ranges. Variable importance analysis indicated a shift in the primary contributing variables from initial motor evoked potential parameters in Model 1 to 1-month clinical functional measures in Model 2. Conclusions: Models integrating 1-month reassessment data showed a tendency toward improved discriminative performance compared to those relying solely on initial data. However, as this study was based on a limited sample from a single institution and instability was observed in certain models, external validation using larger, multicenter cohorts is necessary before generalizing these findings.
Yoo Jin Choo, Min Cheol Chang, Jisu Shin· Journal of Clinical Medicine· 0 citations
Background Diabetic foot ulcers (DFUs) remain the leading cause of non-traumatic lower-extremity amputation (LEA). Most published risk-stratification tools require specialized testing and are difficult to apply in routine practice. We sought to build a simple nomogram for amputation risk using only routinely available clinical indicators. Methods This single-center retrospective cohort study included 508 hospitalized patients with type 2 diabetes-related DFUs treated between January 2019 and June 2025. The cohort was randomly split into training (n=357, 84 amputations) and validation (n=151, 35 amputations) sets at a 7:3 ratio with stratification on outcome. The outcome was lower-extremity amputation (minor or major) during the index hospitalization. Twenty-four baseline variables were screened for selection in the training set (84 events) using two independent methods: LASSO regression with 10-fold cross-validation (at λ.1se) and 1000 bootstrap stepwise logistic regression replicates (selection frequency >80%). Variables retained by both methods were entered into the final multivariable logistic regression model, from which a nomogram was constructed. Discrimination was assessed by the area under the receiver operating characteristic curve (AUC); calibration was evaluated using calibration plots, the Hosmer–Lemeshow test, and the calibration slope; and clinical utility was assessed by decision curve analysis (DCA). Internal validity was further examined by bootstrap resampling with optimism correction. Results Four variables were retained by both LASSO and bootstrap screening: serum albumin (protective), platelet count, smoking history, and hypertension (predictors of higher risk). The nomogram achieved an AUC of 0.788 (95% CI 0.735–0.840) in the training cohort and 0.755 (95% CI 0.671–0.839) in the validation cohort; the wide validation confidence interval reflects the limited number of validation events (n = 35). On bootstrap internal validation (1000 resamples), the optimism-corrected AUC was 0.776. Calibration was acceptable in both cohorts (Hosmer–Lemeshow P = 0.794; calibration slope 0.944 [bootstrap-corrected]), and DCA suggested potential clinical utility. Conclusion A four-variable nomogram based on serum albumin, platelet count, smoking history, and hypertension estimated in-hospital LEA risk in hospitalised patients with type 2 diabetes-related DFUs with moderate discrimination and acceptable calibration, and decision curve analysis suggested potential clinical utility. Because every predictor is available from routine clinical history and standard laboratory testing, the model may support early in-hospital risk stratification rather than treatment decisions. As this was a single-centre study with internal validation only, external multi-centre validation is required before routine clinical use.
Haipeng Zhang, Jian Guo, Yongfang Ma et al.· Diabetes, Metabolic Syndrome...· 0 citations
Background: Peripheral artery disease (PAD) is clinically heterogeneous, complicating identification, treatment selection, and prediction of therapeutic benefit. Gait biomechanics provide high-resolution, objective digital biomarkers that can be used in routine care. Methods: We built a computational framework for three tasks: (1) distinguish controls versus PAD; (2) predict clinician-selected treatment (nonoperative vs. revascularization); and (3) forecast post-treatment ground-reaction force (GRF) from baseline GRF features. The dataset included 55 unique patients with PAD and 42 controls. Standard classifiers and regressors were evaluated, and performance was measured by accuracy, balanced accuracy, Matthew’s correlation coefficient (MCC), and mean absolute error. Results: A single GRF feature (propulsive peak) achieved 90.9% accuracy, 86.7% balanced accuracy, and 0.868 MCC for distinguishing PAD from controls and predicting treatment, outperforming patient-reported outcomes and kinematics. For outcome forecasting, a Random Forest model predicted changes in the propulsive peak with 80% accuracy and low error. Conclusions: GRF-based models accurately identify PAD, anticipate this clinical team’s treatment selection, and forecast functional recovery within this single-center cohort; generalization to other practice patterns and settings remains untested. The approach relies on minimal features and modest computation and is readily generalized to other medical conditions that affect gait. Because it is compatible with wearable sensors, it is well suited to real-time, remote assessment and can advance personalized, data-driven decision support in vascular care.
Ali Al Ramini, Farahnaz Fallahtafti, M. A. Takallou et al.· Applied Sciences· 0 citations
Molecular markers emerged as the strongest predictors, whereas conventional clinical variables showed limited value, and the value of integrating molecular-genetic biomarkers with ML for personalized risk stratification and preventive care in reconstructive surgery is highlighted.
L. Sydorchuk, R. Gumennyi, Miroslav Škoda et al.· Computation· 1 citation