Skip to content
Conference

MedQwen-VGR1: Multimodal Vision-Guided Reasoning Model for Med-VQA, Temporal Diagnosis and Drug Interaction Analysis in Medical Imaging and Pathology

Jul 2026 · 2026 6th International Conference on Electrical, Computer and Energy Technologies (ICECET) · pp. 1-6 · 0 citations · 16 references

Abstract

MedQwen-VGR1 is a novel multimodal visionlanguage model (VLM) trained for medical visual question answering (Med-VQA), longitudinal temporal diagnosis, and drug interaction analysis across radiology and pathology imagery. The model undergoes a multi-stage training pipeline comprising: (1) Continuous Domain-Adaptive Pretraining on heterogeneous medical image-text corpora; (2) Supervised FineTuning on expert-annotated multimodal datasets spanning sequential imaging conversations; (3) Human Preference Alignment via Direct Preference Optimization (DPO) and Group Relative Preference Optimization (GRPO) applied to Vision-Guided Chain-of-Thought (CoT) Reasoning trajectories; and (4) Smart Memory Module integration enables Cross-Visit Context Retention and Comparative Reasoning on sequential or temporal pathologies. MedQwen-VGR1 processes sequential medical images and textual queries to generate stepwise, evidence-grounded rationales correlating visual features across timepoints, predict adverse drug interactions from images/metadata, and output calibrated reliability scores. Evaluations on PathVQA, SLAKE, and VQA-RAD demonstrate superior performance over LLaVA-Med (+14.7% accuracy) and BioViL-T (+12.3% temporal reasoning), achieving 82.4% VQA accuracy and 78.6% change detection precision, enabling multi-turn clinical dialogues and automated clinical reports. The framework enables dynamic, multi-turn clinical dialogues with automated medical report synthesis, addressing critical gaps in temporal reasoning, explainability, and clinical safety alignment for realworld AI and safety enhanced clinical decision support systems.

View source