Low-Rank Adaptation Provides a Parameter-Efficient Trade-Off for Low-Data Medical Imaging
Abstract
Adapting large pretrained Vision Transformers (ViTs) to medical imaging is challenging because labeled clinical data are limited and full fine-tuning is computationally expensive. This paper presents a multi-seed, multi-modality study of parameter-efficient fine-tuning (PEFT) for low-data binary medical image classification. Using a pretrained ViT-Base backbone, we compare linear probing, LoRA with ranks 4, 8, and 16, DoRA, AdaLoRA, and full fine-tuning on BreastMNIST ultrasound and PneumoniaMNIST chest X-ray under a matched low-data protocol. Each dataset-condition pair is repeated over 10 random seeds, and results are evaluated using discrimination, calibration, and Holm-corrected paired significance tests. Full fine-tuning achieves the highest test ROC-AUC on both datasets, reaching 0.908 on BreastMNIST and 0.966 on PneumoniaMNIST, but requires updating 85.8 million parameters. Among PEFT methods, LoRA and DoRA are statistically comparable, while increasing the LoRA rank beyond 8 gives no significant improvement. LoRA at rank 8 retains 90.3% of full fine-tuning’s BreastMNIST AUC and 98.4% of its PneumoniaMNIST AUC while training only 0.35% of the full fine-tuning parameters. Calibration results also show that PEFT methods can produce more reliable probabilities than full fine-tuning despite lower AUC. These findings help guide PEFT method and rank selection when adapting ViTs to low-data medical imaging tasks.