Skip to content
Review

Medical question answering: A comprehensive multimodal and LLM-driven survey.

Jul 2026 · Computer Methods and Programs in Biomedicine · Vol 285, pp. 109541 · 0 citations · 129 references
Medicine

Abstract

Medical Question Answering (MQA) has emerged as a critical artificial intelligence (AI) capability for supporting clinicians, researchers, and the general public with timely and evidence-based responses to medical queries. Recent advances in natural language processing (NLP), computer vision, and large language models (LLMs) have expanded MQA from text-only systems to multimodal frameworks. This survey aims to provide a comprehensive and structured review of MQA systems, covering both text and image-based approaches. We present a systematic review of MQA literature, including applications, datasets, and modeling paradigms. We introduce a unified taxonomy categorizing MQA systems into scientific, clinical, consumer, and examination-oriented tasks. We also analyze representative datasets for text-based and vision-based question answering, focusing on data sources, annotation strategies, task formulations, and evaluation protocols. Furthermore, we review methodological developments ranging from classical and transformer-based models to multimodal vision-language systems and LLM-driven approaches. The analysis highlights a rapid evolution of MQA systems toward multimodal and LLM-based frameworks, particularly in medical visual question answering. Existing datasets and models demonstrate strong progress but also reveal limitations in generalization, reasoning, and real-world clinical applicability. Key challenges remain, including reliability, hallucination, explainability, fairness, and clinical safety. This survey identifies open research directions such as improved data quality, knowledge-grounded reasoning, trustworthy evaluation, and real-world deployment. The study provides a comprehensive reference and roadmap for developing reliable and clinically applicable MQA systems.

View source

Similar papers

Open access Jul 2026

Enhancing medical Q&A systems with multimodal knowledge graphs and dual-layer attention mechanisms

This study develops a text-based intent recognition model with a dual-layer attention architecture, in which a global contextual attention module is introduced to capture long-range semantic dependencies and improve multi-label classification performance.

Guoqiang Qiu, Qingni Yuan, Yi Wang et al. · 0 citations
Review Open access Jul 2026

Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

A dual-view approach that connects clinical practice with computational methods is presented, establishing a five-level competency scheme following Miller’s Pyramid and linking deductive, inductive, and abductive reasoning patterns to common medical goals and tasks.

Qi Peng, Jiatong Li, Sirui Huang et al. · 5 citations
Open access Aug 2026

Retrieval-augmented generation for medical question answering: a multi-metric performance evaluation

The proposed framework offers a practical and scalable approach to mitigating hallucinations without requiring task-specific fine-tuning, highlighting the potential of retrieval-augmented approaches for trustworthy artificial intelligence (AI)-assisted healthcare applications.

Yunus Kökver · 0 citations
Open access Jul 2026

The potential of LLMs in generating questions and answers with EHRs

Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.

Yunqi Zhu, Wen Tang, Huayu Yang et al. · 1 citation
Review Open access Jul 2026

Tutorial: guidance on the use of large language models for medical research

This entry-level tutorial aims to equip healthcare professionals with the tools necessary to effectively integrate LLMs into clinical practice, ensuring that these powerful technologies are applied in a safe, reliable, and impactful manner.

Qiao Jin, Nicholas Wan, Robert Leaman et al. · 1 citation