Preprint
Aug 2026
Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment
A framework that strengthens clinical alignment by separating and enhancing training-time alignment and inference-time alignment is proposed and experiments show that complementary visual representations with a multi-encoder design and concept-level auxiliary learning help preserve clinically meaningful information.
Yunseo Lee, Hyun Jun Kim, Heeseung Shin et al.
· 0 citations