Improving Multimodal Skin Disease Classification via Feature-Space Augmentation with Bayesian Semantic Data Augmentation
Abstract
Automated skin disease classification from clinical photographs faces distinct challenges relative to dermoscopy, including greater illumination variation, class imbalance, and domain shift. We propose a multimodal framework combining a Swin Transformer backbone with a structured metadata encoder and a Bayesian Semantic Data Augmentation (BSDA) module that perturbs the fused image-metadata embedding in feature space rather than at the pixel level. Training uses a three-stage progressive fine-tuning strategy with focal loss. On SkinDisNet, the primary configuration (Swin-S + Metadata + BSDA) achieves 94.73% accuracy and a weighted F1-score of 94.58%, outperforming the multimodal baseline without BSDA by 1.17 percentage points in accuracy. On PAD-UFES-20, the best BSDA-augmented variant reaches 85.69% accuracy, indicating that the strategy remains competitive under a distinct clinical benchmark. Ablation studies confirm that metadata fusion and BSDA provide complementary benefits.