Aug 2026· Abdominal Radiology· 0 citations· 27 references
Medicine
TL;DR
The developed model achieves high sensitivity and precise automated gallstone segmentation on CT images and achieves overall sensitivities of 97.2%, 97.5%, 95.2%, 89.2%, and 98.2% across the training, validation, internal test, hold-out, and AMOS datasets.
Abstract
Purpose
To develop a deep learning model for the automated detection and measurement of gallstones on non-contrast CT images.
Methods
A total of 3,231 CT scans were retrospectively enrolled as an internal cohort, while 753 CT scans from the public AMOS (Abdominal Multi-Organ Segmentation) dataset were employed as an independent external test set. Paired MR imaging served as the reference standard for the internal cohort (including development and hold-out datasets). A three-stage 3D V-Net convolutional neural network was developed for automated gallstone detection. The first stage performed coarse localization of the gallbladder, followed by refined 3D segmentation in the second stage to generate precise anatomical masks. In the final stage, gallstones were detected within the segmented gallbladder volume. Gallstone detection was categorized based on the spatial overlap (Dice similarity coefficient, DSC > 0) between reference labels and predictions.
Results
For gallbladder segmentation, the DSCs were 0.992, 0.989, and 0.990 across training, validation, and internal test sets. Gallstone detection achieved mean DSCs of 0.794, 0.742, and 0.759 across the subgroups. For gallstone detection and segmentation, the model achieved overall sensitivities of 97.2%, 97.5%, 95.2%, 89.2%, and 98.2% across the training, validation, internal test, hold-out, and AMOS datasets, respectively. Subgroup analysis showed median DSCs of 0.835-0.897 for high-density stones and 0.792-0.854 for large stones.
Conclusion
The developed model achieves high sensitivity and precise automated gallstone segmentation on CT images.
To investigate the feasibility of employing deep learning models for automated segmentation and classification of adrenal incidental abnormalities on low-dose CT images.
Four distinct CT cohorts were retrospectively collected for deep learning models development (cohort A,
n
= 2574; cohort B,
n
= 1205), internal evaluation (cohort C,
n
= 3681), and external evaluation (cohort D,
n
= 779). Two experienced uroradiologists independently reviewed the CT images and labeled the adrenal glands as normal or abnormal based on predefined criteria encompassing both density and morphological abnormalities, with any discrepancies resolved through consultation. The model development cohorts were divided into a training set, a validation set, and a test set. Deep learning models for segmentation and classification were trained and evaluated on internal and external sets, with the dice similarity coefficient (DSC), area under precision–recall curves (AUPRC), and area under receiver operating characteristic curves (AUROC) as evaluation metrics. Adrenal descriptions from radiology reports were extracted to compare with the model’s performance.
For adrenal gland segmentation, the DSC values for the test set, internal validation cohort, and external validation cohort were 0.839 (IQR: 0.783–0.871), 0.870 (IQR: 0.819–0.902), and 0.799 (IQR: 0.729–0.849), respectively. For adrenal gland classification, the AI model achieved AUPRC values of 0.913, 0.753, and 0.927 in the test set, internal validation cohort, and external validation cohort, respectively, outperforming routine radiology reporting (AUPRC: 0.809, 0.708, 0.591; all
P
< 0.05). Corresponding AUROC values were 0.956, 0.942, and 0.977 for the AI model, which also outperformed routine radiology reporting (AUROC: 0.889, 0.705, 0.551; all
P
< 0.05).
The deep learning models showed promise in automated adrenal segmentation and classification, highlighting AI’s potential to improve detection of adrenal abnormalities in LDCT scans.
This study has been registered on ClinicalTrials.gov on August 25, 2025, with the unique identifier NCT07198152.
Kexin Wang, He Wang, Shiwei Chen et al.· BMC Medical Imaging· 0 citations
DL enabled accurate CECT-based identification of AP in this retrospective multicenter cohort, with performance maintained in an independent external dataset, and showed promising performance for CECT-based acute pancreatitis detection.
Oleksandra Seidel, M. Theis, Sebastian Nowak et al.· European Radiology Experimen...· 0 citations
Manual evaluation of non-contrast CT scans (NCTS) for detecting subdural hematoma (SDH) is time consuming, potentially inaccurate, and subjective to the expert analyzing them. In recent years, two deep learning (DL) algorithms have been popularly studied in this respect, namely convolutional neural networks (CNN) and U-Net architectures, the latter being a specialized type of CNN. We performed the first meta-analysis comparing various DL models for SDH detection. MEDLINE, Cochrane, Scopus, and Embase databases were searched from inception through December 2025. Studies evaluating ML model performance on an independent test dataset were included. The main outcome measures were sensitivity, specificity, diagnostic odds ratio (DOR), accuracy, and precision of CNN, U-Net, and hybrid DL models. Univariate meta-regression analyses were performed. 30 testing datasets incorporating 67,266 NCTS were included. U-Net demonstrated significantly higher sensitivity (0.916;p = 0.04) and precision (0.983;p = 0.001) while high specificity, DOR, and accuracy values were consistently observed across all DL techniques. Internal testing (p = 0.05) was a borderline significant predictor of high specificity while recent publication year (p < 0.001), U-Net architecture (p = 0.035), and 3D models (p = 0.022) emerged as significant moderators of high precision. The U-Net architecture was also a borderline significant predictor of high DOR (p = 0.049). While this single arm meta-analysis depicts potential superiority of U-Net models with respect to sensitivity and precision, these findings are based off only 4 pooled U-Net datasets in comparison to the 22 pooled for CNN architectures. Future well-powered studies evaluating the U-Net model are necessary to ensure a fair comparison of U-Net architectures to other DL designs before reaching to any definitive conclusions in this respect.
S. Dumasia, W. Ahmed, Michelle Edavettal et al.· Neurosurgical review· 0 citations
The field of radiology is experiencing a surge in demand due to advances in medical imaging, particularly in techniques such as magnetic resonance imaging and computed tomography (CT). However, the interpretation of these scans relies heavily on the availability of experts, which is challenging in resource-limited regions. Recent advances in artificial intelligence and deep learning offer promising solutions by assisting radiologists in image interpretation and diagnosis. This study focuses on validating DeepCTE3D (Deep Convolutional Neural Network for Computed Tomography Extraction 3D), a deep learning–based model based on 3D architecture for segmenting and quantifying intracranial volume (ICV) and lateral ventricular volume (LVV) in CT scans. The model’s performance was evaluated using a real-world dataset comprising diverse patient demographics and various scanner models, including normal and pathological scans. The evaluation process involved developing a streamlined pipeline to generate ground-truth results and comparing them to the model’s outputs. DeepCTE3D achieved high similarity scores for both ICV and LVV. Secondary analyses revealed differences in LVV and ICV between patient sexes and scanner models, although these differences were not clinically significant. This study highlights the potential of DeepCTE3D in enhancing clinical triage and advancing neuroimaging applications, especially in scenarios where MRI is not feasible.
B. Pinto, Tayran Milá Mendes Olegário, Pedro Vinicius Alves Silva et al.· Scientific Reports· 0 citations
This study presents an AI-driven computed tomography (CT) diagnostic system for the automated detection and classification of renal abnormalities in the Bangladeshi population. Renal abnormalities such as kidney cysts, stones, and tumors require timely and accurate diagnosis to reduce complications and improve treatment outcomes. In Bangladesh, the increasing burden of kidney-related diseases has created a strong need for efficient and intelligent diagnostic support systems. To assess the effectiveness of deep learning in multiclass renal abnormality diagnosis, three advanced convolutional neural network architectures were implemented and compared: Xception, VGG16, and ResNet152V2. The experimental results demonstrated outstanding classification performance across all models, with Xception achieving the highest accuracy of 99.84%, followed by ResNet152V2 at 99.76%, and VGG16 at 99.40%. Among the evaluated approaches, Xception showed the best overall performance, indicating its strong capability for reliable renal abnormality classification from CT images. The proposed system has significant potential to assist radiologists and healthcare professionals by providing fast, accurate, and automated diagnostic support. This work highlights the promise of artificial intelligence in medical imaging and contributes to the advancement of intelligent diagnostic solutions for kidney disease detection in Bangladesh.
Mithila Yeasmin Mitu· American Journal of Smart Te...· 0 citations
OBJECTIVES
To develop and evaluate automated deep learning (DL) segmentation of acute neck abscesses on MRI and to assess agreement between automated and manual quantitative measurements relevant to severity assessment.
MATERIALS AND METHODS
In 226 patients with surgically confirmed neck abscesses from a single-center emergency MRI database, a DL segmentation model (nnU-Net v2) was trained using two input configurations: (1) post-contrast T1-weighted images (T1C) only and (2) a three-channel fusion of T1C, T2-weighted fat saturated (T2FS), and apparent diffusion coefficient (ADC) maps, with T2FS and ADC rigidly registered to T1C space. Performance was evaluated using five-fold cross-validation. Spatial overlap (Dice similarity coefficient [DSC]), volumetric agreement (intraclass correlation coefficient [ICC]), and quantitative feature agreement were compared between automated and manual segmentations.
RESULTS
The fusion model achieved a mean DSC of 0.828 ± 0.124 across all 226 cases in five-fold cross-validation. The T1C-only model achieved a mean DSC of 0.813 ± 0.153 and was non-inferior to the fusion model (p = 0.005). Volumetric agreement was excellent for both models (fusion ICC = 0.970, T1C-only ICC = 0.933). Quantitative features showed good to excellent agreement.
CONCLUSION
Automated DL segmentation of neck abscesses on multiparametric MRI is feasible and yields volumetric and quantitative measurements in good to excellent agreement with manual segmentation. The T1C-only model was non-inferior to the three-channel fusion model, offering a parsimonious single-sequence alternative that avoids registration-related confounds.
V. Viertonen, Aapo Sirén, Julius Reima et al.· European Journal of Radiolog...· 0 citations