MedFuse: dual-stream fusion of convolutional and vision transformer-based features for enhanced medical image classification
Abstract
Medical image classification is fundamental to computer-aided diagnosis. Limited labeled samples, subtle inter-class differences, and heterogeneous lesion morphology make it difficult for a single representation to capture all relevant cues. This study investigates MedFuse, a simple dual-stream framework that combines convolutional neural network (CNN) features with representations from a pretrained vision foundation model (VFM). The CNN branch is optimized on the target dataset to emphasize task-adaptive local texture and morphology, whereas the VFM branch is kept frozen to provide a stable global semantic representation. The two pooled feature vectors are concatenated and passed to a common classification head. Accordingly, MedFuse is positioned as a systematic and reproducible fusion baseline. The framework is evaluated across all 12 two-dimensional MedMNISTV2 subsets using multiple CNN-VFM backbone pairs and is compared with alternative weight- and attention-based fusion schemes. Simple late feature concatenation was competitive with, and often more stable than, more parameterized fusion modules across the evaluated settings. Moreover, increasing the size of the VFM backbone did not consistently improve classification performance. These findings demonstrate that frozen foundation-model representations can complement trainable convolutional features in small-scale medical image classification, while also showing that increased fusion complexity or larger VFM backbones does not necessarily translate into better performance. MedFuse therefore provides a simple and reproducible baseline for systematically studying CNN-VFM feature complementarity in medical image classification.