Sep 2026· IEEE Transactions on Medical Imaging· Vol PP, pp. 1-1· 0 citations
Medicine
TL;DR
Experimental results show that PRIMA consistently improves generation quality, question-answering accuracy and expert-rated clinical interpretability compared with existing RAG baselines, demonstrating the effectiveness of clinically grounded prototype-guided retrieval for multimodal knowledge assistance in interventional radiology.
Abstract
Interventional radiology (IR) requires joint reasoning over procedural images and domain-specific clinical knowledge. Existing medical retrieval-augmented generation (RAG) methods are mainly text-oriented or designed for general medical vision-language tasks, and therefore remain limited in retrieving fine-grained visual-textual evidence for IR scenarios. To address this limitation, we present Prototype-guided Retrieval for Interventional Medical Assistance (PRIMA), a multimodal RAG framework that jointly leverages multimodal imaging and clinical text to support IR decision-making. PRIMA constructs a multimodal IR knowledge index through anatomy-aware visual-textual alignment and modality-preserving representation learning. It then introduces domain-informed prototype learning to organize IR concepts, enabling prototype-guided retrieval that re-ranks evidence using both query similarity and prototype affinity. We conduct comprehensive evaluations on literature-curated and clinically collected IR datasets. Experimental results show that PRIMA consistently improves generation quality, question-answering accuracy and expert-rated clinical interpretability compared with existing RAG baselines. These findings demonstrate the effectiveness of clinically grounded prototype-guided retrieval for multimodal knowledge assistance in interventional radiology. Related resources are available at https://github.com/StonHamA/PRIMA.
ClinAlign—a memory-based retrieval framework aligned with clinical workflow, drawing inspiration from clinical diagnostic workflows is proposed, which constructs a disease-aware visual memory bank and introduces Classification-Guided Prompt Augmentation (CGPA), where disease state predictions are converted into structu...
Lihong Qiao, Shi-Yi Gao, Yu-Cheng Shu et al.· Proceedings of the Thirty-Fi...· 0 citations
RG-ICL is introduced, a training-free inference framework that retrieves query-aligned demonstrations using frozen encoders, without task-specific parameter updates, and indicates that for medical image interpretation, curating which reference cases an MLLM sees is a practical alternative to retraining it.
Min-Da Zhao, Fang-Yu Hu, Yan Luo et al.· 0 citations
A novel multimodal RAG framework tailored for MedVQA is proposed, which leverages multimodal data, including medical images, reports, and generated captions, to provide more accurate clinical answers, and introduces a training paradigm that uses captions as auxiliary supervision, enhancing cross-modal alignment via con...
Mai A. Shaaban, M. Zarei, Adnan Khan et al.· 0 citations
Experiments on BUS-CoT and IU X-ray datasets demonstrate consistent improvements in diagnostic accuracy, concept consistency, and report quality over strong general-purpose and medical MLLMs, indicating that concept-grounded reasoning better aligns generation with clinical decision processes.
Xin-Yue Xu, Hong-Bin Lin, Juan-Gui Xu et al.· 0 citations
Automatic radiology report generation has become an active research area due to its potential to reduce radiologist workload and standardize reporting quality. However, state-of- the-art systems still suffer from hallucinated findings, limited clinical reasoning, and a lack of calibrated uncertainty estimates, all of w...
Lokesh P, Kamaleshwaran K, Naveenraj M et al.· 2026 International Conferenc...· 0 citations
Through hierarchical gated co-attention, this approach dynamically aligns image and text representations, addressing the limitations of static fusion and providing a foundation for interpretable, real-time retrieval systems that can accelerate and improve clinical decision support.