Skip to content
Review

Advancements in Medical Information Fusion: A Review of Deep Learning Models and Their Applications

Aug 2026 · WIREs Data Mining and Knowledge Discovery · Vol 16 · 0 citations · 59 references

TL;DR

This review systematically explores deep learning‐driven advancements, focusing on their theoretical foundations, key technological breakthroughs and typical application scenarios, and systematically categorizes fusion architectures into data‐level, feature‐level, and decision‐level hierarchies across diverse modalities.

Abstract

Recent advancements in deep learning have transformed medical image processing, particularly in integrating multimodal data and optimizing complex tasks, positioning fusion models as the leading paradigm in contemporary research. Historically viewed as mere pixel‐level blending, this field has fundamentally evolved into a broader paradigm of Medical Information Fusion. This review systematically explores deep learning‐driven advancements, focusing on their theoretical foundations, key technological breakthroughs and typical application scenarios. First, we systematically categorize fusion architectures into data‐level, feature‐level, and decision‐level hierarchies across diverse modalities. Second, we address key technical challenges and methodological breakthroughs, focusing on advanced solutions for domain shift and semantic alignment, handling data scarcity via missing modality completion, and task‐specific optimization through multi‐task learning. Furthermore, we critically examine the clinical transitional potential of these models in key clinical scenarios, including advanced tumor diagnosis, smart surgery, personalized therapy, and brain function analysis. By synthesizing these insights, this work aims to provide a comprehensive understanding of the current landscape and future directions of medical information fusion, paving the way for advancements in precision medicine and improved healthcare outcomes.

View source

Similar papers

Review Open access Aug 2026

A Comprehensive Review Tracing the Evolution of Volumetric Medical Imaging Analysis from Classic CNNs to Emerging AI-Agents

Volumetric medical imaging has redefined modern healthcare, enabling precise diagnosis, prognosis, and treatment planning. During the past decade, the field has undergone a paradigm shift from classical deep learning architectures to multimodal, agent-driven AI systems capable of uncovering rich volumetric biomarkers and utilizing heterogeneous data for predictive and generative modeling. Existing surveys are fragmented, focusing on specific models or tasks instead of offering a unified view of volumetric learning evolution. This study traces the evolution from classical models (Convolutional Neural Networks, Recurrent Neural Networks, and transformers) to generative approaches (Variational Autoencoders, Generative Adversarial Networks, and diffusion models) and finally to foundation models and AI-agents that enable advanced reasoning and adaptive clinical workflows. As the reported performance varies substantially across datasets, imaging modalities, and evaluation protocols, this review emphasizes methodological evolution, representative innovations, and practical implications rather than direct numerical ranking of competing architectures. For each paradigm, we critically assess methodological innovations, strengths, limitations, and comparative performance in segmentation, classification, detection, reconstruction, and report generation. Beyond synthesizing progress, we identify persistent challenges, including data scarcity, generalization between institutions, and clinical trustworthiness, and outline emerging frontiers in multimodal fusion, explainable AI, and human–AI collaboration. This review provides a unified framework for understanding the evolution of volumetric medical imaging and offers actionable insights for researchers, clinicians, and industry practitioners, contributing to the development of reliable, interpretable, and clinically deployable next-generation medical AI systems. To support further research, we provide a GitHub repository that includes popular 3D medical imaging datasets with recent 3D models in our shared GitHub repository (https://github.com/Owais-CodeHub/3D-Medical-Imaging-Review).

Muhammad Owais, Muhammad Zubair, Daniya Najiha Abdul Kareem et al. · 8 citations · ⚡1
Review Open access Jul 2026

A comprehensive review of efficient deep learning for clinical medical imaging deployment

This review provides a comprehensive synthesis of efficient and lightweight deep learning architectures specifically tailored for the medical domain, and examines key model compression strategies and their efficacy in maintaining diagnostic performance while reducing hardware requirements.

C. M. Nguyen, Truong-Son Hy · 0 citations
Open access 2024

Deep Neural Frameworks for Integrating Multimodal Healthcare Data

Multimodal data integration is gaining traction in medical image analysis, enabling the use of diverse data sources to improve downstream tasks. Deep Learning approaches have proliferated, employing generic architectures and a data-driven paradigm. While initial efforts have yielded positive results, they lack inherent adaptation to the peculiarities of medical multimodality. Bridging representations across signal pairs and aligning disparate modalities provide more robust performance. Specifically, representation learning, explicitly learning transferable feature extraction models, has emerged as an important research avenue. Contrastive learning and visual-language pre-training provide methods to learn joint embedding spaces. The proposed multimodal evaluation setup examines several public datasets, offering a well-designed statistical analysis framework and research-practice reproducibility. Baseline models explore early-and late-fusion scenarios for multimodal emotion recognition and sickness prediction from facial expression. Initial results indicate representative power and proper data alignment as crucial elements. As multimodality gains momentum in Deep Learning research, bridging modalities and demonstrating clear real-world applications pave the way for impactful contributions.

Unknown authors · 0 citations
Review Jul 2026

Exploring Deep Transfer Learning for Medical Image Processing and Analysis: A Comprehensive Analysis across Modalities.

Empirical evidence from recent studies demonstrates that fine-tuning and network-based DTL strategies, including federated learning, consistently enhance diagnostic accuracy, robustness, and generalization across multiple medical imaging modalities, particularly in data-limited clinical scenarios.

M. A. S. Banu, A. Dhavapandiammal, K. Palanisamy · 0 citations
Book Open access Jul 2026

Advancing Multimodal Fusion on Heterogeneous Medical Data with Hybrid Geometry Attention

Multimodal fusion learning (MFL) (a framework to jointly learn from heterogeneous data sources) has shown great potential in various fields such as Medicine, Science, and Engineering. It is extremely desirable in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges. First, they struggle to capture complex cross-modal interactions effectively, which in turn limits performance improvements. Second, they incur high computational costs, restricting their applicability in resource-constrained healthcare AI applications. Finally, they are often designed and evaluated for narrow, fixed modality configurations (e.g., imaging-only, or specific pairs such as image and omics), which limits evidence of their adaptability and generalizability to broader collections of heterogeneous medical modalities. To address these challenges, we propose a novel MFL framework – Cascaded Unified Representation Learning for Efficient Fusion Network (CURE) – a lightweight and scalable framework that progressively integrates various modalities through a novel efficient Hybrid Geometry Aware Fusion layer (HyFuse), where each HyFuse layer is sequentially learned for each modality, making the framework adaptable and generalizable. Within HyFuse, an efficient residual convolution module captures rich multi-scale features to ensure cost-effective learning, while a hybrid-space aware attention mixer learns coarse-to-fine structural cues to better preserve cross-modal relationships. Complementary learnable late-fusion and shared-information refinement modules are then employed to learn robust, modality-order-invariant shared features, which in turn yields consistent performance improvements. Extensive evaluations on 16 public datasets show that CURE outperforms leading multimodal fusion methods (e.g., DRIFA-Net and HEALNet), boosting performance by up to ≈ 3.97% and lowering computational costs by up to ≈ 87.8%, ensuring more effective and reliable predictions.

Joy Dhar, M. Pandey, Nayyar Zaidi et al. · 0 citations