Vision-language models (VLMs) have significantly advanced open-vocabulary image understanding by learning aligned representations from large-scale image-text datasets. Despite their zero-shot generalization capabilities, adapting these foundation models to specific application domains remains challenging. Full fine-tuning is often infeasible due to computational costs and the risk of overfitting when labeled data are limited. Parameter-efficient fine-tuning (PEFT) approaches promise to address this issue by updating only a small set of parameters while keeping the pre-trained encoders largely frozen. However, existing PEFT strategies often exhibit insufficient spatial awareness, rendering them suboptimal for medical imaging, where subtle visual differences can lead to distinct clinical diagnoses. To address this challenge, in this study we propose a novel PEFT technique that leverages natural-image pretrained DINOv3’s attention maps to enforce spatial alignment.
M. A. Aydın, Furkan Genc¸, Efe ¨Ozdilek et al.· Signal Processing and Commun...· 0 citations
Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) are fundamental imaging modalities that provide complementary information for clinical assessment. However, CT acquisition may not always be preferred or available due to additional cost, workflow burden, and exposure to ionizing radiation. Therefore, reliably synthesizing a corresponding CT image from an available MRI scan is an important medical imaging problem. In this work, we present a conditional latent diffusion framework for MRI-to-CT synthesis. To improve the reliability of stochastic generation, the proposed framework incorporates a modular output steering mechanism that favors candidates better matched to target-domain characteristics. In this way, the diversity of diffusion-based synthesis is preserved while output consistency and realism are improved. Experimental results indicate that the proposed approach yields improved quantitative and qualitative performance over a baseline latent diffusion model. The proposed method achieved 26.49±2.19 dB PSNR and 87.89±2.59% SSIM, outperforming the baseline latent diffusion model.
Efe Özdilek, Fuat Arslan, Boran Ismet Macun et al.· Signal Processing and Commun...· 0 citations