Aug 2026· Diseases of the esophagus· Vol 39· 0 citations
TL;DR
Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset, and this training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.
Abstract
Benign Disease: New Technologies
Hiatal hernia diagnosis using barium swallow studies often requires multiple image acquisitions to visualize the esophagogastric junction adequately. Repeated acquisitions may increase radiation exposure and barium ingestion. Interpretation is observer-dependent. The aim of this study was to train an artificial intelligence model for detecting and classifying hiatal hernias.
A limited dataset of 70 anonymized barium swallow images centered on the esophagogastric junction was retrospectively analyzed and classified as hiatal hernia (n=45) or normal (n=25). Images were standardized using region-of-interest cropping, autocontrast adjustment, resizing to 512×512 pixels, and data augmentation to mitigate small-sample limitations. The dataset was divided into training (70%), validation (15%), and testing (15%) subsets.
Three pretrained convolutional neural networks (ResNet18, DenseNet121, EfficientNetB0) were fine-tuned using transfer learning. A structured grid search (72 trials) optimized learning rate, dropout, weight decay, and training epochs under identical cross-validation conditions. The primary evaluation metric was mean area under the curve (AUC). Secondary metrics included F1-score, accuracy, sensitivity, and specificity. Grad-CAM visualization was applied to assess anatomical regions influencing predictions.
Despite the limited dataset and use of focused images, all architectures demonstrated strong internal performance. DenseNet121 achieved mean AUC and F1 values of 1.00 in cross-validation, while EfficientNetB0 showed the strongest out-of-fold performance (AUC 0.9733; F1 0.9556). ResNet18 achieved mean cross-validation AUC 0.9956 and F1 0.9882.
Under a unified grid-search comparison framework, ResNet18 demonstrated the most consistent global performance (mean AUC 0.7860; F1 0.7662; accuracy 0.7571; sensitivity 0.6215; specificity 1.0000), supporting its selection as the final model. Grad-CAM confirmed consistent attention to the esophagogastric junction region. The prototype web application generated automated binary classification with confidence scoring.
Deep learning models can support detection and classification of hiatal hernias using focused AP barium swallow images, even when trained on a limited dataset. This training approach may enable automated identification and classification while potentially reducing radiation exposure and barium ingestion during diagnostic studies.
Timely identification of children with ileocolic intussusception likely to fail air-enema reduction is critical to avoid delays and bowel perforation. However, even expert sonographers show inter-observer variability. We developed and prospectively validated a Vision Transformer (ViT) deep learning system to predict reduction failure from static B-mode ultrasound images. This multicenter bidirectional cohort study included 5602 children (4-60 months) who underwent air-enema reduction at 14 Chinese tertiary hospitals (retrospective cohort: 2019-2024). After data augmentation, 10,151 images (8122 training, 2029 validation) were used to train a ViT model for binary classification ("success" vs. "failure"). External validation was performed on a prospective cohort of 190 patients (March-June 2025), with three junior and three senior sonographers independently predicting outcomes. The study was approved by the Ethics Committee of Yijishan Hospital of Wannan Medical University (approval No. 2025-04) and registered with ChiCTR2500098673. The model achieved high internal performance (failure: accuracy 0.880, precision 0.969; success: accuracy 0.970, precision 0.898). In the prospective cohort, the ViT model achieved 93.7% overall accuracy, significantly higher than senior (74.7%) and junior (60.7%) sonographers (p < 0.05). This study innovatively applies ViT to assess pediatric ileocolic intussusception severity, providing an objective, accurate tool to support clinical decision-making and reduce treatment risks.
Jie Liu, Yue Wang, Danping Zeng et al.· npj Digital Medicine· 0 citations
The proposed framework establishes a reliable and lightweight baseline for automated gastrointestinal disease detection and demonstrates that ConvNeXt-Tiny effectively captures disease-relevant visual patterns in endoscopic images while maintaining consistent performance across varying training conditions.
Muhammad Faqih, O. Q. Aziz, Ajib Hanani· Jurnal Ilmu Komputer dan Inf...· 0 citations
An advanced system capable of automatically detecting complicated appendicitis from ultrasound images was developed and was explained with gradient-weighted class activation mapping (Grad-CAM), which creates a heatmap of the regions responsible for the model's prediction of the infected areas.
Fahad Ahammed, Omar Faruq Shikdar, Navid Zaman et al.· 0 citations
Background: Deep-learning systems for endoscopic image classification are commonly evaluated with random data splits, which may overestimate performance under acquisition shift; uncertainty-aware selective prediction may improve reliability by allowing a model to abstain on uncertain cases. Methods: We evaluated binary abnormality detection in two endoscopic imaging domains: laryngeal contact-endoscopy narrow-band imaging (CE-NBI; 210 patients, patient-level) and gastrointestinal endoscopy (HyperKvasir; 6746 images, image-level, predominantly white-light). ImageNet-pretrained ResNet-50 deep ensembles were assessed under a resolution-defined acquisition-shift stress test; for the gastrointestinal data, a random stratified split was additionally used as an in-distribution reference. We evaluated discrimination, calibration, decision-curve analysis, and entropy-based selective prediction. Results: In the gastrointestinal dataset, random-split evaluation produced high performance (AUROC 0.994, 95% CI 0.990–0.996; AUPRC 0.991). Under resolution shift on the same data, performance fell to AUROC 0.723 (95% CI 0.710–0.736; AUPRC 0.690; sensitivity 0.417; specificity 0.874); the two intervals do not overlap. Selective prediction improved reliability among retained cases: under resolution shift, accuracy rose from 0.683 at full coverage to 0.852 (95% CI 0.835–0.871) at 25% coverage (balanced accuracy 0.646 → 0.798). Predictive entropy was significantly higher for incorrect than for correct predictions in both regimes (Mann–Whitney p = 7.1 × 10−77 with rank-biserial |r| = 0.30 under shift). In the laryngeal cohort, no statistically significant differences were detected among four architectures (ROC-AUC 0.844–0.901; all pairwise DeLong p > 0.05). Conclusions: Random-split evaluation substantially overestimated performance relative to a resolution-defined acquisition-shift stress test, and entropy-based selective prediction improved reliability by identifying a high-confidence subset for automated prediction while deferring the remainder to human review. Target-domain recalibration substantially restores calibration under shift (ECE 0.172 → 0.036 with temperature scaling; → 0.017 with isotonic regression) but does not recover discrimination; selective prediction is complementary, mitigating residual confident-wrong predictions. An encoder-transfer experiment showed asymmetric cross-domain utility; features learned on the larger gastrointestinal cohort transferred to the laryngeal cohort (AUROC 0.80 vs. in-domain 0.89), whereas the reverse direction did not transfer (0.53 vs. 0.72). Prospective multi-center validation remains required before clinical deployment. All code, fold definitions, random seeds, and a reproducible protocol are publicly released.
Behnam Kiani Kalejahi, S. Khan, Murodbek Akhrorov et al.· Biomedicines· 0 citations
While the results are promising for a novel application domain, the model's failure on clinically critical minority classes (Bilateral Blockage, Bilateral Patency) means it is not yet suitable for unsupervised clinical use.
Nasreen Jawaid, I. Brohi, Najma Imtiaz Ali et al.· Scientific Reports· 0 citations
The feasibility of developing a deep learning model for the detection of intussusception using a smaller dataset of POCUS images is demonstrated and fine-tuning models were best adapted to the screening nature of POCUS images.
A. Thyagachandran, Brian Lefchak, H. Murthy et al.· Frontiers in Radiology· 0 citations