Difference Visual Question Answering in Medical Imaging Based on Multimodal Large Models and Enhanced by Prior Region-level Difference Descriptions
Abstract
Difference Visual Question Answering (Diff-VQA) in medical imaging automatically compares patient images across time points to support assessment of lesion progression and treatment efficacy. However, pixel-level matching is unreliable due to non-rigid deformations, view shifts, and acquisition noise, while existing models often rely on synthetic labels and lack effective integration of local and global information. To address these challenges, we propose a multimodal large-model framework that adopts a progressive “local semantic modeling–global difference reasoning” strategy. Key anatomical regions in chest X-rays are localized via object detection and aligned with VinDr-CXR annotations to construct region–disease mappings, transforming misalignment into semantic difference analysis. A dynamic sampling strategy further generates clinically meaningful image pairs with fine-grained difference labels. Finally, a multimodal large model fuses local features with global context to support single-image QA, dual-image disease description, and global difference reasoning. Experiments on the MIMIC-Diff-VQA dataset demonstrate state-of-the-art accuracy in single-image QA and substantial improvements in Diff-VQA tasks over mainstream medical large models. In the single-image QA tasks, our model improves accuracy from 52.5% to 64.1% (22.2% relative improvement), and in the Diff-VQA tasks, the CIDEr score increases from 1.027 to 1.379 (34.3% relative improvement). These results highlight the framework’s potential to enhance diagnostic accuracy and strengthen clinical decision support in radiology practice.