Aug 2026· Frontiers in Digital Health· Vol 8· 0 citations· 77 references
TL;DR
It is argued that LLM-based CDSS are not supported for routine autonomous use and require clinician supervision, and set out a research agenda centred on validation, governance and human-in-the-loop deployment.
Abstract
Deep learning (DL) and generative artificial intelligence (generative AI) are changing how medical data are analysed and used at the point of care. Our evidence base comprises 80 sources 37 screened studies and 43 landmark primary studies, architectural papers and clinical-AI reporting standards published between 2014 and 2026, across two related domains: medical image analysis and clinical decision support systems (CDSS). We trace the evolution from convolutional neural networks (CNNs) to U-Net and encoder–decoder networks, then to Vision Transformers (ViTs), generative adversarial networks (GANs), diffusion models, large language models (LLMs) and retrieval-augmented generation (RAG). Each model family is compared across nine dimensions: input modality, task, data requirements, validation level, interpretability, failure modes, clinical readiness, regulatory considerations and human-oversight need. Supervised DL reaches clinically useful performance across CT, MRI and pathology on well-scoped detection, segmentation and classification tasks, though results are task-, dataset- and site-dependent and prospective evidence is limited. To address data scarcity, generative models can produce synthetic images or cross-modality translations, but may amplify hidden dataset biases and generate anatomically incorrect images. LLM-based CDSS show promise for guideline-concordant reasoning and medication-safety checks, yet still face hallucination, calibration and regulatory uncertainty. For each technology we provide a deployment-readiness map, review the reporting standards needed for credible clinical evaluation (CONSORT-AI, SPIRIT-AI, TRIPOD + AI, CLAIM, DECIDE-AI, STARD-AI, PROBAST + AI, FUTURE-AI) and set out a research agenda centred on validation, governance and human-in-the-loop deployment. Unlike prior surveys, which treat imaging AI or clinical LLMs separately, we assess both, add an explicit evidence-quality appraisal and link each technology to the reporting standards. On current evidence, we argue that LLM-based CDSS are not supported for routine autonomous use and require clinician supervision. This is not a PRISMA-style systematic review but a structured critical narrative review built on a curated, transparently reported evidence base.
Recent advances in generative artificial intelligence (AI) have shown significant promise for brain magnetic resonance imaging (MRI), enabling applications such as image synthesis, modality translation, reconstruction, super-resolution, segmentation, anomaly detection, and disease identification. This PRISMA-ScR-guided scoping review provides a structured synthesis of recent peer-reviewed studies on generative AI for brain MRI analysis published between January 2024 and March 2026. A total of 43 studies meeting predefined inclusion criteria were analyzed. We review major generative architectures, including variational autoencoders (VAEs), generative adversarial networks (GANs), diffusion models, and transformer-based generative models, and summarize their applications across key neuroimaging tasks. We also provide an overview of the publicly available datasets commonly used for model development and evaluation. Beyond reporting performance, this review critically examines the current evidence with respect to reproducibility, external validation, data availability, evaluation validity, data leakage, hallucination and safety risks, and barriers to clinical translation. Although many studies report promising results on retrospective benchmark datasets, external validation, prospective evaluation, reader studies, and clinically oriented assessments remain relatively uncommon. Challenges related to generalization, dataset heterogeneity, computational requirements, privacy, and regulatory considerations continue to limit real-world deployment. Overall, the reviewed literature demonstrates that generative AI has substantial potential to improve brain MRI analysis through realistic data generation, enhanced image quality, and more informative feature representations. However, the current evidence primarily supports technical feasibility and methodological advances rather than established clinical utility. We conclude by identifying key research gaps and future research directions toward more robust, interpretable, reproducible, and clinically translatable generative AI frameworks for brain MRI analysis.
Ahmed Kammoun, M. Akhloufi· Information· 0 citations
PURPOSE
Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias.
METHODS
Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative.
RESULTS
Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain.
CONCLUSION
Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.
Md Mazharul Islam, Abrar Mohammed Tanzim Alam, Md Sadikur Rahman Rony et al.· International Journal of Med...· 0 citations
Volumetric medical imaging has redefined modern healthcare, enabling precise diagnosis, prognosis, and treatment planning. During the past decade, the field has undergone a paradigm shift from classical deep learning architectures to multimodal, agent-driven AI systems capable of uncovering rich volumetric biomarkers and utilizing heterogeneous data for predictive and generative modeling. Existing surveys are fragmented, focusing on specific models or tasks instead of offering a unified view of volumetric learning evolution. This study traces the evolution from classical models (Convolutional Neural Networks, Recurrent Neural Networks, and transformers) to generative approaches (Variational Autoencoders, Generative Adversarial Networks, and diffusion models) and finally to foundation models and AI-agents that enable advanced reasoning and adaptive clinical workflows. As the reported performance varies substantially across datasets, imaging modalities, and evaluation protocols, this review emphasizes methodological evolution, representative innovations, and practical implications rather than direct numerical ranking of competing architectures. For each paradigm, we critically assess methodological innovations, strengths, limitations, and comparative performance in segmentation, classification, detection, reconstruction, and report generation. Beyond synthesizing progress, we identify persistent challenges, including data scarcity, generalization between institutions, and clinical trustworthiness, and outline emerging frontiers in multimodal fusion, explainable AI, and human–AI collaboration. This review provides a unified framework for understanding the evolution of volumetric medical imaging and offers actionable insights for researchers, clinicians, and industry practitioners, contributing to the development of reliable, interpretable, and clinically deployable next-generation medical AI systems. To support further research, we provide a GitHub repository that includes popular 3D medical imaging datasets with recent 3D models in our shared GitHub repository (https://github.com/Owais-CodeHub/3D-Medical-Imaging-Review).
Muhammad Owais, Muhammad Zubair, Daniya Najiha Abdul Kareem et al.· Archives of Computational Me...· 8 citations· ⚡1
A definitive taxonomy of the medical VLM landscape is provided, tracing the evolution from early Contrastive Alignment and Generative MLLMs to the cutting-edge frontiers of Dense Pixel-Grounding, Sparse Mixture-of-Experts (MoE), and Reasoning-Incentivized (RL) architectures.
Taha Razzaq, Murtaza Taj, Asim Iqbal· Journal of Biomedical Inform...· 0 citations
ABSTRACT
The adoption of whole-slide imaging is establishing a new paradigm in digital pathology. However, the translation of artificial intelligence (AI) from research to clinical practice faces significant hurdles, largely due to a misalignment between algorithmic advances and the practical demands of pathological diagnosis and prognosis. In this review, we propose a dual-perspective framework to systematically bridge this gap by linking core clinical tasks with cutting-edge deep learning methodologies. We present a comprehensive overview of the field from 2020 to 2025, analyzing how architectures such as convolutional neural networks, vision transformers, and graph neural networks are being adapted for diagnostic classification, tissue segmentation, and prognostic prediction. A key contribution is our novel algorithm-clinical task mapping framework, which offers practical guidance for selecting and designing AI solutions tailored to specific clinical goals. We also highlight emerging trends that minimize reliance on costly annotations-including weakly supervised and self-supervised learning-as well as advances in predicting immunohistochemistry results directly from hematoxylin and eosin-stained slides. Finally, we address critical challenges related to model interpretability, regulatory approval, and multicenter generalization, and outline a future pathway focused on developing integrated, trustworthy, and equitable AI systems that enhance, rather than replace, the expertise of pathologists.
Yun-qiu Gao, Teng Ma, Lisha Li et al.· Chinese Medical Journal· 0 citations