Skip to content
Conference

AI Driven Disease-Specific Adaptive Multimodal Medical Image Fusion with CNN

Jul 2026 · 2026 4th International Conference on Sustainable Computing and Smart Systems (ICSCSS) · pp. 1676-1681 · 0 citations · 17 references

Abstract

Despite the utilization of multimodal medical imaging as supplementary anatomical and functional data crucial for precise illness diagnosis, the appropriate integration of multimodal pictures has been challenging due to discrepancies in resolution, contrast, and disease-specific imaging characteristics. Most established techniques for image fusion are modality-specific, depend on manually crafted features, and lack the adaptability to encompass a wide variety of pathological presentations, hence constraining their clinical applicability. This paper proposes an AI-driven, disease-specific, adaptive multimodal medical image fusion utilizing Convolutional Neural Networks (CNNs) to tackle these challenges. This strategy is proposed to be an end-to-end trained modality-specific and pathology-aware feature capable of generating adaptive fusion to increase clinically significant areas, excluding structural characteristics. The qualitative and quantitative assessments reveal that the experimental results of multimodal medical image datasets indicate that the proposed method surpasses both traditional and sophisticated deep learning-based fusion techniques. The performance metrics of PSNR, SSIM, entropy, and mutual information, which indicate enhanced fusion quality, contrast, and diagnostic clarity, substantiate the effectiveness of the proposed framework in facilitating disease-oriented clinical decision-making and computer-aided diagnostics systems.

View source

Similar papers

Conference Jul 2026

Multi-Modal Medical Image Fusion Using Hybrid CNN-Transformer Models for Early Detection of Chronic Diseases

Chronic disease early and accurate detection is a major healthcare issue nowadays, and most importantly, there is the rising prevalence or use of heterogeneous medical image data such as CT, MRI, X-ray, and retinal scans. Conventional models of deep learning such as CNNs perform well on spatial aspects of feature extraction but not generally on long-term relations and overall context. In this paper, we introduce a new hybrid deep learning network combining Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) to carry out multi-modal medical image fusion to assist in the diagnostic process with better results. The CNN branch gives fine-grain local representation and the Transformer module will predict the global connections between modalities. Empirical tests on publicly available data show that the proposed model is better than single CNN and Transformer models in using the datasets and there are massive increments in precise score, recall, and F1-score on chronic disease diagnosis in the initial stages. The field of research shows how hybrid architecture can be used to combine complementary information about various scanning modalities, which can become the direction of AI-aided decision-making in prevention.

N.R Azhakeshwari, S. Christy, S. Saranya et al. · 0 citations
Open access Aug 2026

Advancements in  Medical Image Processing: Integrating Physics With Convolutional Neural Networks (CNN) for Enhanced Diagnostic

Background: Medical imaging has innovatively changed the healthcare set-up through the facilitation of early and non-invasive analysis. Despite all these, the developing complexity of imaging information needs qualitative diagnosis tools. Although deep learning, especially with the use CNNs, has indicated possible, challenges such as data dependency and interpretation may hamper its clinical acceptability.  Objective: The major focus of this study may add physical limitations to CNNs to improve consistency and healthcare organizational adoption of deep learning frameworks. The main scope of this work may include the growth of a physics-Informed Convolutional Neural Network (Pi-CNN) for sort of brain tumour categorization from MRI images and followed by its testing framework. Methodology: Scientists employed baseline CNN model and Pi-CNN with the same parameters to MRI brain tumor exercise procedures. With spatial consonance limitation being integrated into the Pi-CNN it can enable the medical image forecast through the method of anatomy. Homogenous metric evaluations were carried out on their experimentation plan by several researchers who applied the two different model types. Results: At the testing stage, the Baseline CNN may set out to attain a slightly better precision rate at 84% than Pi-CNN which achieved at 80.0%. Information detection precision and generalization abilities were better in the Pi-CNN which affirmed that the status is the good choice for the purpose of medical application. Specific Contribution: The precise input delivers a replication to Physics-Informed CNN model, which maintains a precise results whereas tempting to enhance the quality of the results and clinical trust phases. The base framework will enable deep learning approaches at numerous levels to combine clinical perfection with outcomes which can enhance diagnostic imaging performance.   Conclusion: Medical professionals applying limitations for CNNs identified enhancement in more efficient models which can enhance clinical utility and consistency. Future AI model innovation may involve improved pre-existing purview of knowledge of biological pedigree for physics in connection with enhanced health record translation because of the presumed combination and it will assist physicians in their investigation procedures.

Unknown authors · 0 citations
Review Open access Aug 2026

A trustworthy cross-domain AI framework for fundus disease classification using hybrid CNN fusion and supervised domain adaptation

Deep learning–based systems for fundus disease classification often achieve impressive accuracy on internal datasets, yet their performance degrades markedly when applied to real-world clinical data. This limitation is primarily caused by domain shift arising from variations in imaging devices, illumination conditions, and population characteristics, which remains a key barrier to integrating these technologies into reliable clinical decision-support systems. To address this challenge, we propose FusionEye-Net, a hybrid deep learning framework that performs feature-level fusion of EfficientNet-B3 and ResNet-50 for four-class fundus image classification, namely cataract, diabetic retinopathy, glaucoma, and normal retina. A dedicated preprocessing pipeline—including circular fundus cropping, illumination normalization, contrast-limited adaptive histogram equalization (CLAHE), and adaptive gamma correction—was employed to reduce inter-device variability and standardize retinal appearance. FusionEye-Net was first trained on a curated internal dataset of approximately 4000 images, achieving an internal test accuracy of 99.24%. However, evaluation on an independent external dataset comprising 4640 images revealed a significant performance drop to 72.05%, highlighting the severity of cross-domain variability. To mitigate this degradation, a targeted supervised domain adaptation strategy was applied using a balanced subset of 2000 external images (500 per class). This adaptation improved external performance to 91.27% accuracy with a macro F1-score of 0.9139, reducing misclassifications by more than two-thirds. Model predictions and visual explanations were reviewed by a board-certified retina specialist to ensure clinical plausibility. Explainability analysis using Gradient-weighted Class Activation Mapping (Grad-CAM) and region-of-interest (ROI) contour mapping demonstrated that the adapted model consistently focuses on medically relevant structures, such as lens opacity in cataract, microaneurysms in diabetic retinopathy, and optic disc cupping in glaucoma. In summary, the results demonstrate that FusionEye-Net combines strong internal performance with substantially improved cross-domain generalization through lightweight adaptation, underscoring the critical role of external validation and domain-aware fine-tuning in formulating trustworthy, adaptive AI-assisted decision-support tools that can safely augment automated ophthalmic screening workflows.

Ali M. Duhaim, A. M. Al-Bakry · 0 citations
Jul 2026

High Performance Hybrid CNN CBAM Framework for High Sensitivity Heart Disease Classification

Cardiovascular diseases remain a paramount global health crisis, necessitating early and precise diagnostic interventions. While medical imaging is the clinical standard, manual interpretation is highly susceptible to visual fatigue and inter-observer variability. This study proposes a novel, highly robust Computer-Aided Diagnosis (CAD) framework that overcomes the spatial and textural limitations of standalone Convolutional Neural Networks (CNN) in heart disease image classification. A Feature-Level Ensemble (Hybrid) architecture was created by putting together the deep semantic features of ResNet50V2, the spatial boundaries of VGG16, and the parameter efficiency of EfficientNetV2B3. To directly deal with the loss of features caused by anatomical background noise, a Convolutional Block Attention Module (CBAM) was added to the EfficientNet pathway. This gave the network two-dimensional (channel and spatial) visual attention. To guarantee a thorough and impartial assessment, a complete restructuring of a dataset comprising 5,977 images was undertaken using an 80:10:10 stratified split, thereby eliminating the accuracy paradox resulting from class imbalance. The proposed Hybrid CBAM model significantly outperforms standalone baselines, with a peak accuracy of 94.00%. For clinical use, it was very important that the attention-guided ensemble had a Recall (sensitivity) of 0.94 for finding pathological cases and a Negative Precision of 0.96. This study definitively demonstrates that the integration of multi-model feature extraction with focused visual attention mechanisms yields a highly sensitive, reliable, and non-invasive automated screening instrument for the early detection of cardiovascular disease.

Giant Prakoso Amukti Wibowo, Slamet Riyadi, A. Dewi et al. · 0 citations
Open access Jul 2026

Multimodal Brain Tumour Classification and Segmentation Using a Dual-Attention Swin-UNet with Quantum-Inspired Optimisation for MRI-Based Diagnostics

One of the most significant and challenging challenges in medical image processing is the use of magnetic resonance imaging (MRI) to diagnose brain cancers. Significant tumour heterogeneity, irregular morphologies, blurred boundaries, and differences between imaging modalities are the root causes of the problem. Early and precise disease detection is critical for better treatment planning and patient outcomes. Automated brain tumour analysis has greatly improved because of deep learning, particularly convolutional neural networks (CNNs). However, with high-resolution medical images, traditional CNN-based models frequently struggle to identify long-range contextual linkages. This limitation may make it more difficult to obtain precise segmentation and reduce the overall reliability of the diagnosis. Transformer-based designs have recently shown great potential in visualising global relationships. However, because they are hard to train, require a lot of memory, and are rather complex to utilise, their practical application in medical imaging is challenging. We provide a Multimodal Brain Tumour Classification and Segmentation framework based on a Dual-Attention Swin-UNet architecture, enhanced with a Quantum-Inspired Optimisation (QIO) technique, to overcome these limitations. The proposed system successfully integrates fine-grained spatial localisation and global contextual awareness by combining a U-Net-style decoder with an encoder based on Swin Transformer. Dual attention mechanisms spatial and channel attention are added to enhance feature representation and facilitate visualisation of tumour edges, thereby further accelerating feature learning. These attention modules filter away irrelevant background noise, allowing the network to concentrate on clinically significant areas. Additionally, we introduce a quantum-inspired optimisation+ technique to reduce training oscillations and enhance convergence stability, which are frequently observed in transformer-based medical imaging models. A pre-processed version of the BraTS 2020 dataset and additional brain MRI images recorded in HDF5 format are used to train and evaluate the framework. This enables more effective data handling and large-scale training. Due to computational power constraints, GPU acceleration is used for full-scale model training. A real-time, user-friendly web-based diagnostic interface demonstrates a lightweight and deployable inference pipeline. The experimental outcomes demonstrate the strength and clinical practicability of the proposed method, which offers superior tumour segmentation accuracy and reliable classification. The possible development of scalable and intelligent clinical decision support systems for brain tumour diagnosis through the integration of transformer-based architectures, the geography of attention processes, and quantum-inspired optimisation strategies is identified in the study.

Anirban Mondal, V. E. Jesi · 0 citations