DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset, and showed promising performance for CECT-based acute pancreatitis detection.
Abstract
We developed and evaluated a deep learning (DL) model for image-based detection of acute pancreatitis (AP) on abdominal contrast-enhanced CT (CECT). A total of 552 patients from two university centers (January 2010–January 2026) were included. The internal dataset comprised 207 patients with clinically and radiologically confirmed AP (499 scans) and 250 control patients with suspected AP (368 scans). An independent external validation cohort included 95 patients. Convolutional neural network–based models were trained using monophasic and biphasic CECT data. The final model was evaluated on a 20% patient-level hold-out test set from the internal cohort and on the external cohort. Performance was assessed using the F1 score and area under the receiver operating characteristic curve (AUROC). A single-input multiphase model incorporating arterial and portal venous phase scans achieved the best performance, with ensembling applied for biphasic studies. On the internal hold-out test set (n = 116), the model achieved an F1 score of 0.83 (95% confidence interval 0.75–0.89) and an AUROC of 0.89 (0.82–0.95). Performance remained robust on external validation (n = 95), with an AUROC of 0.99 (0.96–1.00) and an F1 score of 0.92 (0.86–0.97). DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset. Prospective validation using broader and independently adjudicated clinical populations remains necessary. The model showed promising performance for CECT-based acute pancreatitis detection but was not designed or tested as a triage system. Diagnostic uncertainty in acute pancreatitis often arises from nonspecific abdominal symptoms and inter-reader variability in CECT interpretation. The DL model achieved high internal accuracy (AUROC 0.89) and maintained robust performance in an independent external validation cohort (AUROC 0.99). Diagnostic uncertainty in acute pancreatitis often arises from nonspecific abdominal symptoms and inter-reader variability in CECT interpretation. The DL model achieved high internal accuracy (AUROC 0.89) and maintained robust performance in an independent external validation cohort (AUROC 0.99).
To investigate the feasibility of employing deep learning models for automated segmentation and classification of adrenal incidental abnormalities on low-dose CT images.
Four distinct CT cohorts were retrospectively collected for deep learning models development (cohort A,
n
= 2574; cohort B,
n
= 1205), internal evaluation (cohort C,
n
= 3681), and external evaluation (cohort D,
n
= 779). Two experienced uroradiologists independently reviewed the CT images and labeled the adrenal glands as normal or abnormal based on predefined criteria encompassing both density and morphological abnormalities, with any discrepancies resolved through consultation. The model development cohorts were divided into a training set, a validation set, and a test set. Deep learning models for segmentation and classification were trained and evaluated on internal and external sets, with the dice similarity coefficient (DSC), area under precision–recall curves (AUPRC), and area under receiver operating characteristic curves (AUROC) as evaluation metrics. Adrenal descriptions from radiology reports were extracted to compare with the model’s performance.
For adrenal gland segmentation, the DSC values for the test set, internal validation cohort, and external validation cohort were 0.839 (IQR: 0.783–0.871), 0.870 (IQR: 0.819–0.902), and 0.799 (IQR: 0.729–0.849), respectively. For adrenal gland classification, the AI model achieved AUPRC values of 0.913, 0.753, and 0.927 in the test set, internal validation cohort, and external validation cohort, respectively, outperforming routine radiology reporting (AUPRC: 0.809, 0.708, 0.591; all
P
< 0.05). Corresponding AUROC values were 0.956, 0.942, and 0.977 for the AI model, which also outperformed routine radiology reporting (AUROC: 0.889, 0.705, 0.551; all
P
< 0.05).
The deep learning models showed promise in automated adrenal segmentation and classification, highlighting AI’s potential to improve detection of adrenal abnormalities in LDCT scans.
This study has been registered on ClinicalTrials.gov on August 25, 2025, with the unique identifier NCT07198152.
Kexin Wang, He Wang, Shiwei Chen et al.· BMC Medical Imaging· 0 citations
The purpose of this study is to develop and retrospectively validate a deep learning system for the classification and anatomically interpretable localization of six acute abdominal emergencies on CT using multi-window Hounsfield Unit (HU) encoding. A publicly available national teleradiology dataset of 1274 patients (42,922 bounding box annotations) was used for training and internal validation (896/189/189 patient-level split). Each CT slice was encoded into three diagnostic HU windows (soft tissue, bone/stone, angio/liver). A YOLOv11-Large model with a stride-4 (P2) head was trained at 1280 × 1280 . Localization was evaluated using the clinical nine-region abdominal grid, and specificity was assessed in 80 target-negative patients. External validation used a radiologist-adjudicated 280-patient Stanford Merlin cohort (US), with model weights and thresholds applied without modification. Internal macro AUROC was 0.941, and macro F1 was 76.1%. Nine-region localization accuracy was 99.5% among detected cases and 90.9% including missed detections. Specificity in the target-negative cohort was 86.2%. On the external Stanford Merlin cohort, macro AUROC was 0.879 with all six classes ≥ 0.80 ; AAA reached F1 0.889 at frozen thresholds; macro F1 was 0.545, increasing to 0.648 after recalibration. A multi-window CT detection model classified six acute abdominal emergencies with high discrimination on internal testing and preserved moderate-to-high discrimination on an external Stanford Merlin cohort. Region-level evaluation using the nine-region abdominal grid provided a clinically interpretable localization endpoint complementary to conventional detection metrics. However, variable class-level F1, reduced external operating-point performance, and retrospective design indicate that multisite prospective validation and site-specific threshold calibration are required before clinical deployment.
Hasan Mete Erdoğan, Ural Koç· Journal of imaging informati...· 0 citations
Acute pancreatitis (AP) incidence is rising globally. Current scoring systems lack sensitivity for early organ failure (OF) prediction and suffer from interobserver variability. This study aimed to develop and validate an artificial intelligence (AI)-driven model for fully automated early prediction of OF in AP using multiphase computed tomography (CT) imaging.
This multicenter study included 2746 AP patients from two tertiary hospitals (2011–2024). Patients were split into training (
n
=1820), validation (
n
=456), and test cohorts (
n
=470). An nnMamba-based segmentation model delineated pancreatic/peripancreatic regions on CT. An organ failure risk assessment with CT and learning engine (ORACLE) model integrated deep learning radiomics (severe organ failure–deep learning radiomics [SOF-DLR] score from 57 optimal features) with clinical variables. The primary outcome was OF (Modified Marshall Score≥2).
OF occurred in 8.7% (
n
=240). The ORACLE model achieved the receiver operating characteristic curves (area under the curve [AUCs]) of 0.85 (training), 0.89 (validation), and 0.81 (test), outperforming Modified CT Severity Index (M-CTSI) (AUC 0.68–0.74) and clinical models (AUC 0.67–0.71; DeLong’s
P
<0.001). The overall negative predictive value for the entire cohort (
n
=2746) was 97.2%. High-risk patients (
P
>0.700; 1.4% of cohort) had 92.1% OF incidence. The model provided a median early warning time of 3.5 hours (mean 9.17 h) before clinical OF onset, with 55% of cases predicted ≥3 h in advance.
This AI-based tool enables accurate, automated OF prediction 3.5 h before clinical manifestation, facilitating risk-stratified management. Its generalizability is confirmed in multicenter validation.
Yifei Guo, Chengwei Chen, Tiegong Wang et al.· Journal of Pancreatology· 0 citations
BACKGROUND
The accurate identification of children with refractory Mycoplasma pneumoniae pneumonia (RMPP) remains challenging. This study aimed to develop a transformer-based model utilizing clinically indicated chest computed tomography (CT) to stratify pediatric RMPP risk at a critical decision point.
METHODS
Non-contrast chest CT data from a multicenter retrospective cohort of 1224 pediatric patients with Mycoplasma pneumoniae pneumonia who underwent clinically indicated CT were used to develop a transformer-based deep learning framework (trans-DLF). The primary cohort comprised training (n = 506), validation (n = 140), and internal testing (n = 139) cohorts, with two independent external cohorts (n = 331 and n = 108) used to evaluate generalizability. Model performance was assessed by the area under the receiver operating characteristic curve (AUC) and compared against a three-dimensional convolutional neural network (3D-CNN), a clinical model, and a multimodal nomogram. Interpretability was examined using gradient-weighted class activation mapping (Grad-CAM).
RESULTS
The median age was 6.83 years (interquartile range, 5.0-8.6 years), and 609 (49.8%) were male. The trans-DLF demonstrated strong performance across all cohorts: training (AUC 0.97; 95% confidence interval [CI], 0.96-0.98), validation (0.91; 0.86-0.96), internal testing (0.90; 0.85-0.95), and external testing (0.89; 0.84-0.94 and 0.89; 0.82-0.95). It significantly outperformed the clinical model (p < 0.001), while its AUCs were not significantly different from those of the multimodal nomogram. The model maintained good performance in outpatient settings (AUC 0.87) with good calibration and net clinical benefit. Grad-CAM suggested that predictions were influenced by clinically meaningful features, particularly consolidations.
CONCLUSION
The trans-DLF provides a streamlined and efficient approach to RMPP risk assessment in children who have already undergone clinically indicated chest CT and may support timely, evidence-based decision-making without additional tests.
Zhoumeng Ying, Ge Hu, Jing Li et al.· BMC Medical Imaging· 0 citations
Accurate preoperative discrimination of renal cell carcinoma (RCC) subtypes is critical for treatment stratification. We aimed to develop and validate an automated deep learning system for simultaneous tumor segmentation and histopathological subtyping using multicenter contrast-enhanced CT (CECT) imaging. To this end, we constructed a two-stage system comprising separate segmentation and classification models. The segmentation model was trained on 245 scans from Nanfang Hospital and 210 from the KiTS19 public dataset. The classification model was developed and validated on a total of 750 patients, comprising an internal cohort from Nanfang Hospital (553 patients; 328 training, 112 validation, 113 testing) and two external validation cohorts: one from Beijing Tongren Hospital (n = 111) and another combined from two other centers (n = 86). The model demonstrated strong generalizability for discriminating clear cell RCC, with AUCs of 0.878 (internal validation), 0.892 (internal testing), 0.911 (external set Ⅰ), and 0.892 (external set Ⅱ). The model's computational efficiency reached 0.24 s per file and reduced FLOPs by four times compared to conventional 3D CNNs. This study validates the efficiency and clinical applicability of the YOLOv11 framework for RCC subtyping. Future efforts should integrate prospective data and multimodal imaging to enhance sensitivity for small lesions.