Aug 2026· Frontiers in Artificial Intelligence· Vol 9· 0 citations· 130 references
Medicine
TL;DR
The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks.
Abstract
Transformer-based architectures have become central to medical image analysis, yet their practical value remains difficult to assess because studies vary widely in tasks, datasets, validation protocols, baselines, and reporting quality. This survey critically reviews recent transformer-based, hybrid, foundation, and transformer-alternative models across segmentation, classification, reconstruction, and image registration. A total of 128 studies published between 2021 and 2026 are organized using a task-, modality-, and architecture-aware taxonomy, with reported performance synthesized alongside baseline comparisons, reproducibility, computational cost, and clinical-readiness evidence. The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks. Persistent gaps include non-standardized benchmarks, limited external validation, incomplete code and weight availability, inconsistent efficiency reporting, weak uncertainty analysis, and insufficient clinical evaluation. Progress will require transparent reporting, multicenter validation, and clinically grounded assessment.
The growing integration of Transformer-based architectures into 3D medical image analysis has driven significant advances across segmentation, classification, detection, registration, and reconstruction tasks. However, existing reviews remain fragmented, often focusing on 2D medical image analysis or specific modalities or tasks without providing a comprehensive, structured synthesis of architectural innovations, benchmark performance, and clinical applicability. This systematic review addresses these gaps by following the PRISMA 2020 framework to systematically evaluate 138 peer-reviewed studies published between January 2017 and December 2025, identified across Scopus, Web of Science, and PubMed. We evaluate Transformer-based and hybrid CNN–Transformer architectures across major 3D imaging modalities; MRI, CT, PET, and ultrasound using a structured five-question research framework addressing architectural evolution, benchmark performance, modality-specific trends, methodological rigour, and reproducibility. To quantify methodological progress beyond performance metrics, we introduce an Architectural Innovation Score grounded in a component-level innovation encoding scheme. Our analysis reveals a clear dominance of hybrid CNN–Transformer architectures, with MRI representing the most extensively studied modality. Multimodal imaging achieves the highest normalized mean performance, followed by MRI, CT, and ultrasound. Emerging paradigms, including State Space Models and diffusion-based Transformers, show promise but remain underexplored. Despite strong benchmark results, critical limitations persist, including insufficient external validation, limited code availability, and inconsistent reporting practices.
Wandile Nhlapho, M. Atemkeng, Yusuf Brima et al.· Artificial Intelligence Revi...· 0 citations
Major deep learning architectures, including CNNs, residual networks, UNet, attention-based models, Vision Transformers, and hybrid approaches, along with their clinical applications are summarized and emerging directions such as self-supervised learning, Explainable AI, federated learning, and lightweight models are highlighted as promising approaches for more reliable and accessible medical image analysis.
Lakshmi Sai Anusha Dadi, Pravallika Devi Kommana· International Journal for Re...· 0 citations
This review provides a comprehensive synthesis of efficient and lightweight deep learning architectures specifically tailored for the medical domain, and examines key model compression strategies and their efficacy in maintaining diagnostic performance while reducing hardware requirements.
C. M. Nguyen, Truong-Son Hy· Discover Artificial Intellig...· 0 citations
Clinicians can quickly find previous cases that are visually and semantically similar to a query image using content-based medical image retrieval (CBMIR); however, its practical implementation is hindered by multiple modalities, a lack of labeled data, and vulnerability to real-world artifacts like cropping and blurring. This research presents a hybrid Neural Network-Transformer encoder trained on four complimentary datasets: the COVID-19 Radiography Database, the Kvasir GI endoscopy dataset, and the ImageCLEFmed 2007 and 2009 radiology benchmarks. The model creates compact 128-dimensional embeddings that integrate particular local cues with more general anatomical contexts using a ResNet-50 convolutional backbone, a lightweight transformer head, and a gated fusion module. In order to improve robustness and cross-domain generalization, this study employs data augmentation that mimics real corruptions during training and evaluates retrieval under both clean and intentionally degraded queries. Using a unified methodology based on Precision@K, Recall@K, and mean average precision (mAP), the augmented hybrid model maintains competitive accuracy across all four datasets while producing more retrieval-relevant embeddings and much more consistent performance under degraded queries. This suggests that the model may serve as a promising foundation for future CBMIR research and potential clinical exploration.
Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.
A. Vergara-Richart, Xavier Rafael-Palou, A. Fuster-Matanzo et al.· 0 citations