Skip to content
Review Open access

A survey of transformer-based architectures in medical image analysis: models, applications, and challenges

Aug 2026 · Frontiers in Artificial Intelligence · Vol 9 · 0 citations · 130 references
Medicine

TL;DR

The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks.

Abstract

Transformer-based architectures have become central to medical image analysis, yet their practical value remains difficult to assess because studies vary widely in tasks, datasets, validation protocols, baselines, and reporting quality. This survey critically reviews recent transformer-based, hybrid, foundation, and transformer-alternative models across segmentation, classification, reconstruction, and image registration. A total of 128 studies published between 2021 and 2026 are organized using a task-, modality-, and architecture-aware taxonomy, with reported performance synthesized alongside baseline comparisons, reproducibility, computational cost, and clinical-readiness evidence. The findings indicate that the most convincing gains arise from task-adapted hybrid designs that combine local feature extraction with global context modeling, rather than from an unconditional superiority of transformers over convolutional networks. Persistent gaps include non-standardized benchmarks, limited external validation, incomplete code and weight availability, inconsistent efficiency reporting, weak uncertainty analysis, and insufficient clinical evaluation. Progress will require transparent reporting, multicenter validation, and clinically grounded assessment.

Read PDF

Similar papers

Review Open access Jul 2026

Transformers for 3D medical image analysis: a systematic review of architectural innovations, performance, and clinical applications

The growing integration of Transformer-based architectures into 3D medical image analysis has driven significant advances across segmentation, classification, detection, registration, and reconstruction tasks. However, existing reviews remain fragmented, often focusing on 2D medical image analysis or specific modalities or tasks without providing a comprehensive, structured synthesis of architectural innovations, benchmark performance, and clinical applicability. This systematic review addresses these gaps by following the PRISMA 2020 framework to systematically evaluate 138 peer-reviewed studies published between January 2017 and December 2025, identified across Scopus, Web of Science, and PubMed. We evaluate Transformer-based and hybrid CNN–Transformer architectures across major 3D imaging modalities; MRI, CT, PET, and ultrasound using a structured five-question research framework addressing architectural evolution, benchmark performance, modality-specific trends, methodological rigour, and reproducibility. To quantify methodological progress beyond performance metrics, we introduce an Architectural Innovation Score grounded in a component-level innovation encoding scheme. Our analysis reveals a clear dominance of hybrid CNN–Transformer architectures, with MRI representing the most extensively studied modality. Multimodal imaging achieves the highest normalized mean performance, followed by MRI, CT, and ultrasound. Emerging paradigms, including State Space Models and diffusion-based Transformers, show promise but remain underexplored. Despite strong benchmark results, critical limitations persist, including insufficient external validation, limited code availability, and inconsistent reporting practices.

Wandile Nhlapho, M. Atemkeng, Yusuf Brima et al. · 0 citations
#explainable ai Review Open access Aug 2026

Deep Learning in Medical Imaging: Architectures, Clinical Applications, and Emerging Directions

Major deep learning architectures, including CNNs, residual networks, UNet, attention-based models, Vision Transformers, and hybrid approaches, along with their clinical applications are summarized and emerging directions such as self-supervised learning, Explainable AI, federated learning, and lightweight models are highlighted as promising approaches for more reliable and accessible medical image analysis.

Lakshmi Sai Anusha Dadi, Pravallika Devi Kommana · 0 citations
Review Open access Jul 2026

A comprehensive review of efficient deep learning for clinical medical imaging deployment

This review provides a comprehensive synthesis of efficient and lightweight deep learning architectures specifically tailored for the medical domain, and examines key model compression strategies and their efficacy in maintaining diagnostic performance while reducing hardware requirements.

C. M. Nguyen, Truong-Son Hy · 0 citations
Open access Aug 2026

Hybrid CNN–transformer architecture for content based medical image retrieval

Clinicians can quickly find previous cases that are visually and semantically similar to a query image using content-based medical image retrieval (CBMIR); however, its practical implementation is hindered by multiple modalities, a lack of labeled data, and vulnerability to real-world artifacts like cropping and blurring. This research presents a hybrid Neural Network-Transformer encoder trained on four complimentary datasets: the COVID-19 Radiography Database, the Kvasir GI endoscopy dataset, and the ImageCLEFmed 2007 and 2009 radiology benchmarks. The model creates compact 128-dimensional embeddings that integrate particular local cues with more general anatomical contexts using a ResNet-50 convolutional backbone, a lightweight transformer head, and a gated fusion module. In order to improve robustness and cross-domain generalization, this study employs data augmentation that mimics real corruptions during training and evaluates retrieval under both clean and intentionally degraded queries. Using a unified methodology based on Precision@K, Recall@K, and mean average precision (mAP), the augmented hybrid model maintains competitive accuracy across all four datasets while producing more retrieval-relevant embeddings and much more consistent performance under degraded queries. This suggests that the model may serve as a promising foundation for future CBMIR research and potential clinical exploration.

Namita Bhatt, Alongbar Wary, Ritika Kumari · 0 citations
Review Jul 2026

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.

A. Vergara-Richart, Xavier Rafael-Palou, A. Fuster-Matanzo et al. · 0 citations