Aug 2026· Italian National Conference on Sensors· Vol 26· 0 citations· 138 references
Medicine
TL;DR
This review provides a comprehensive overview of MMIF from theoretical and technical perspectives and presents current challenges and future development trends of MMIF, including robustness, clinical translation, and emerging multimodal learning paradigms.
Abstract
In recent years, multimodal medical image fusion (MMIF) has attracted significant attention due to its ability to integrate complementary information from different medical imaging modalities and provide more comprehensive information for clinical analysis. By combining anatomical information from modalities such as computed tomography (CT) and magnetic resonance imaging (MRI) with functional information from positron emission tomography (PET) and single-photon emission computed tomography (SPECT), MMIF can improve image quality and support subsequent tasks such as disease analysis, lesion detection, image segmentation, and treatment planning. This review provides a comprehensive overview of MMIF from theoretical and technical perspectives. First, commonly used medical imaging modalities and publicly available medical image databases are summarized and compared. Subsequently, the general workflow, fusion levels, and quality requirements of MMIF are introduced. Representative fusion techniques are then systematically reviewed, including spatial-domain methods, transform-domain methods, sparse representation-based methods, deep learning-based methods, hybrid methods, and emerging Mamba-based approaches. In addition, commonly used image fusion quality assessment metrics are analyzed, and the reported quantitative performance of representative MMIF methods is compared and discussed. Finally, current challenges and future development trends of MMIF are presented, including robustness, clinical translation, and emerging multimodal learning paradigms.
Medical image fusion (MIF) is the process of fusing two medical pictures from different modalities into one image. This technology tries to produce a fused output image from two source images that contain more effective and relevant information. This image is used in the healthcare industry, specifically for disease diagnosis. The main challenge is using a single image modality to diagnose diseases accurately. The fused image includes spectral and structural information for the source images to help doctors with disease diagnosis problems. Positron emission tomography (PET), magnetic resonance imaging (MRI), computed tomography (CT), and single photon emission computed tomography (SPECT) are some of the medical imaging modalities. Each modality has its benefits and drawbacks. Researchers have presented different MIF techniques that obtain high fusion results in the MIF field. This paper is a comprehensive survey of multiple state-of-the-art MIF techniques in the spatial and transform domains. It also discusses the main MIF evaluation metrics. Finally, quantitative and qualitative evaluations for some of these techniques are obtained.
Abstract
Purpose To examine how combining magnetic resonance (MR), computed tomography (CT), and positron emission tomography (PET) enhances clinical outcomes in complex or metastatic cancers, focusing on diagnostic performance, staging accuracy, and effects on treatment decisions.
Method Data were extracted and analyzed using a thematic synthesis approach for this literature review. Key variables included study design, sample size, cancer type, imaging modalities used, and reported diagnostic outcomes such as sensitivity, specificity, and accuracy.
Results The integration of multiple imaging modalities enhanced diagnostic performance, improved differentiation between tumor recurrence and posttreatment changes, and increased detection of occult metastatic disease. Advanced imaging approaches, including total-body PET/CT and hybrid PET/MR systems, further improved staging accuracy and allowed for more precise assessment of tumor burden.
Discussion Findings demonstrated that combined imaging modalities possess inherent strengths. MR offered superior soft tissue resolution and high sensitivity, particularly in detecting local tumor recurrence, while PET/CT provided valuable metabolic and whole-body imaging capabilities, improving detection of metastatic disease.
Conclusion Despite these advantages, limitations such as small sample sizes, retrospective designs, and variability in imaging protocols were noted across studies. Overall, the evidence supported the growing role of multi-modal imaging as a critical tool in precision oncology.
Despite the utilization of multimodal medical imaging as supplementary anatomical and functional data crucial for precise illness diagnosis, the appropriate integration of multimodal pictures has been challenging due to discrepancies in resolution, contrast, and disease-specific imaging characteristics. Most established techniques for image fusion are modality-specific, depend on manually crafted features, and lack the adaptability to encompass a wide variety of pathological presentations, hence constraining their clinical applicability. This paper proposes an AI-driven, disease-specific, adaptive multimodal medical image fusion utilizing Convolutional Neural Networks (CNNs) to tackle these challenges. This strategy is proposed to be an end-to-end trained modality-specific and pathology-aware feature capable of generating adaptive fusion to increase clinically significant areas, excluding structural characteristics. The qualitative and quantitative assessments reveal that the experimental results of multimodal medical image datasets indicate that the proposed method surpasses both traditional and sophisticated deep learning-based fusion techniques. The performance metrics of PSNR, SSIM, entropy, and mutual information, which indicate enhanced fusion quality, contrast, and diagnostic clarity, substantiate the effectiveness of the proposed framework in facilitating disease-oriented clinical decision-making and computer-aided diagnostics systems.
K. Jameema, T. Sunitha, Maruturi.Haribabu et al.· 2026 4th International Confe...· 0 citations
Multimodal medical image fusion integrates complementary information from CT, MRI, and PET to support clinical diagnosis and lesion localization. CNN-based methods are constrained by local receptive fields and fail to model long-range dependencies, while Transformer-based methods incur quadratic computational complexity; both paradigms suffer from inadequate inter-modal redundancy suppression, producing blurred edges and structural inconsistency. To address these limitations, Bi-CMFM is proposed, comprising the Bi-Path Residual Fusion module (BPRF), which employs parallel standard and dilated convolutions to preserve local texture while expanding the receptive field, and the Cross-Modal Fusion Module (CMFM), which applies a Cross-modal Feature Enhancement component (CFEM) and a selective scanning mechanism to model long-range dependencies at linear complexity. Experimental results demonstrate that Bi-CMFM achieves EN = 5.1281 and AG = 7.0127 on CT-MRI fusion, PSNR of 12.7624/19.6562 on PET-MRI/SPECT-MRI tasks, and top-ranked EN, SF, AG, SCD, and CC on KAIST, outperforming six representative baselines including DATFuse, DRCM, and FusionMamba.
Jingdong Yang· Poster Volume 0008 The 2026...· 0 citations
Major deep learning architectures, including CNNs, residual networks, UNet, attention-based models, Vision Transformers, and hybrid approaches, along with their clinical applications are summarized and emerging directions such as self-supervised learning, Explainable AI, federated learning, and lightweight models are highlighted as promising approaches for more reliable and accessible medical image analysis.
Lakshmi Sai Anusha Dadi, Pravallika Devi Kommana· International Journal for Re...· 0 citations