Jul 2026· The Lancet Digital Health· Vol 8, pp.
101007
· 2 citations· 23 references
Medicine
TL;DR
Multimodal, Multi-Disease Medical Imaging Foundation Model MerMED-FM has the potential to be a highly adaptable, versatile, cross-specialty foundation model that enables robust interpretation of medical imaging across diverse medical disciplines.
Abstract
Background
Current artificial intelligence (AI) models for medical imaging predominantly focus on a single imaging modality and a single disease. Attempts to create multimodal and multi-disease models have resulted in inconsistent clinical accuracy. Furthermore, training these models typically requires large, well labelled datasets, which are costly and labour intensive to prepare. We aimed to train and evaluate an AI model that can interpret diverse imaging modalities across specialties while maintaining robust performance within each modality.
Methods
We developed Multimodal, Multi-Disease Medical Imaging Foundation Model (MerMED-FM), a multi-specialty model trained using self-supervised learning and a memory module. MerMED-FM was pretrained on publicly sourced, unlabelled medical images from 12 specialties and seven imaging modalities: chest x-rays, CT, ultrasound, histopathology, colour fundus photography (CFP), optical coherence tomography (OCT), and dermatoscopy. After pretraining, the model was fine-tuned, validated, and evaluated for the diagnosis of a range of diseases on 26 public datasets and five private datasets comprising radiology, histopathology, and ophthalmology images. MerMED-FM was compared against a general-domain vision foundation model, various specialist single-modality foundation models, and a multispecialty foundation model. Models were fine-tuned using 10%, 30%, 50%, and 100% of data, with primary comparative analyses conducted using a 10% label fraction. The primary outcome was the area under the receiver operating characteristic curve (AUROC), which was summarised by imaging modality.
Findings
MerMED-FM was trained on around 3·3 million images from 53 publicly available, unlabelled datasets, comprising 713 931 chest x-rays, 292 353 CT slices, 389 885 ultrasound frames, 1 017 712 pathology patches, 333 099 CFP images, 176 719 OCT slices, and 401 059 dermatoscopy images. Strong performance was achieved across all modalities at a label fraction of only 10%, with mean AUROC values of 0·844 for chest x-rays, 0·906 for CT, 0·818 for ultrasound, 0·908 for histopathology, 0·810 for CFP, 0·962 for OCT, and 0·827 for dermatoscopy.
Interpretation
MerMED-FM has the potential to be a highly adaptable, versatile, cross-specialty foundation model that enables robust interpretation of medical imaging across diverse medical disciplines.
Funding
National Medical Research Council, Singapore and the Agency for Science, Technology and Research, Singapore.
An engineering-oriented deployment framework is proposed, integrating modality-driven model selection, structured preprocessing pipelines, multi-level clinical validation, computational feasibility assessment, and explainability, together with a clinical deployment readiness model spanning validation maturity, data div...
Enoch Jacob Dodo, Amos Takai Yayock, Gregory Onwodi et al.· Journal of Science Research...· 0 citations
The project is the medical diagnostic system, powered by AI and based on the Convolutional Neural Networks (CNNs), that automatically identifies abnormalities in medical images with a high precision that is expected to enhance the accuracy of the diagnosis, ease the workload of the medical specialists, and offer scalab...
B. Aravind, Budati, Sheela Koushi et al.· 0 citations
Precision medicine is shifting from a single-biomarker paradigm toward multimodal integration of molecular, imaging, and clinical data, with artificial intelligence (AI) serving as a key enabling technology. Multimodal imaging, computational pathology, spatial omics and liquid biopsy have all become more adept at captu...
Rong Wei, Qi-Ping Zheng· American Journal of Clinical...· 0 citations
Multi-modal learning has demonstrated strong potential in medical applications by integrating heterogeneous data sources such as medical imaging, clinical records, and genomics to improve predictive performance and support clinical decision-making. However, advances in this area are often constrained by two key challen...
Rita Cordeiro Mendes, Maria Rita Verdelho, Carlos Santiago et al.· 0 citations
This review summarizes the deep-learning architectures, fusion strategies, representative applications, and implementation challenges of mpMRI-centered multimodal AI.
M. Fujiwara, Soichiro Yoshida, Imon Banerjee et al.· Abdominal Radiology· 0 citations
Multi-modal learning combining medical images and clinical text is promising for disease diagnosis. However, standard multi-modal training leads to shortcut learning: models exploit the easier modality (e.g., diagnostic cues in text) while neglecting harder-to-learn features (e.g., subtle visual patterns). We propose U...
Zijian Gu, Weikai Lin, Shuang Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.