Skip to content

IndicMedQA: Multimodal Medical Query Analysis in Indian Languages

Jul 2026 · ACM Transactions on Computing for Healthcare · 1 citation · 54 references

TL;DR

This work introduces IndicMedQA, a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues, and creates a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach.

Abstract

In the rapidly advancing field of AI-driven telehealth services, effective medical communication in India remains a significant challenge due to the language barrier, as most of the population is not proficient in English. Additionally, many patients struggle to accurately describe their medical conditions using text alone. As a result, an essential feature of any telehealth service is the ability to supplement textual queries with medical images, enabling doctors to conduct a more careful analysis and provide well-informed diagnoses and treatment recommendations. In this work, we introduce IndicMedQA , a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues. To support this, we create a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach. This dataset facilitates medical understanding for every input query and its associated medical image—the output is a detailed patient summary, symptom analysis, probable conditions, additional findings, and severity assessment with medical precision. Our framework significantly enhances personalized healthcare experiences, ensuring context-aware multimodal understanding of patient needs in Indic languages. Extensive experiments demonstrate that IndicMedQA surpasses all baselines, establishing a new benchmark for Indic AI in healthcare. Disclaimer: This work contains medical images that depict the subject matter of the study, which may be disturbing to some readers.

View source

Similar papers

Open access Jul 2026

NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria

A novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria reveals several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts.

Tobi Olatunji, C. Aka, C. Okocha et al. · 0 citations
Preprint Aug 2026

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

A multilingual medical VQA benchmark over eight languages is constructed, organized into four representative scenarios that isolate the core capabilities medical VQA requires, and a training-free scenario-aware representation engineering method is proposed, leveraging LVLMs's superior English medical VQA capability to steer non-English representations toward their English counterparts at inference time.

Jingbo Wang, Sendong Zhao, Haochun Wang et al. · 0 citations
Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
Preprint Jul 2026

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

A large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital, MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation.

Runhan Shi, Quan Zhou, Yuqian Xu et al. · 0 citations
#natural language process... Preprint Aug 2026

Gaokerena: A Small Persian Medical Language Model Family

Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer grade hardware, and Gaokerena-R, a novel family of compact Persian medical language models optimized for deployment on consumer grade hardware, are presented.

Mehrdad Ghassabi, Hamidreza Baradaran Kashani, Pedram Rostami et al. · 0 citations