Skip to content
Open access

NigBench: A multilingual point-of-care medical query benchmarking study of large language models in Nigeria

Jul 2026 · medRxiv · 0 citations
Medicine

TL;DR

A novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria reveals several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts.

Abstract

In this study, we introduce a novel benchmark comprising over 9,000 real-world, point-of-care, multilingual, and multimodal clinical question-answer pairs sourced from frontline health workers in Nigeria. Using the dataset, we compare local general practitioners to multiple leading open and closed LLMs. Our results reveal several critical insights into the suitability of LLMs as clinical decision support systems in low-resource contexts. The results confirm that performance varies widely by language and input modality (e.g., text vs speech): while models perform best on English text inputs, their accuracy drops significantly for local-language speech. Critically, it is possible to achieve substantial performance gains by transcribing and translating other languages into English before prompting an LLM-- an important insight for non-anglophone product developers. Finally, this benchmark highlights key limitations of SLMs in supporting frontline healthcare in low-resource settings and provides a clear opportunity to track improvements as novel solutions are developed.

Read PDF

Similar papers

Preprint Aug 2026

A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower, and this results indicate that a purpose-built clinical RAG system remains competitive with frontier LLMs on an open benchmark, consistent with corpus specificity as a design variable that improves grounding at some cost to communication polish.

Praveen Reddy, C. Mandke, Suvrankar Datta et al. · 0 citations
Review Open access Jul 2026

Domain-specific versus general large language models: a review and empirical benchmark in real medical texts

It is suggested that domain-adapted encoder models may be preferable for similar structured clinical NER settings, although larger and externally validated benchmarks are needed before generalizing to other languages, clinical corpora, model families, or deployment environments.

L. Elvas, Carolina Carvalho · 0 citations
Jul 2026

IndicMedQA: Multimodal Medical Query Analysis in Indian Languages

This work introduces IndicMedQA, a novel multimodal AI framework that integrates Indic large language models (LLMs) and visual encoders to analyze patient inquiries using both textual and visual cues, and creates a multilingual multimodal medical corpus spanning seven major Indian languages, translated using a semi-automated approach.

Akash Ghosh, Arkadeep Acharya, M. Muhsin et al. · 1 citation
Review Open access Jul 2026

A Large Language Model Leaderboard for Clinical Note Entity Extraction

An LLM leaderboard showing how open-source LLMs perform at entity extraction on unseen clinical notes is developed, showing that large language models are already available that can perform entity extraction well enough to be considered in place of some administrative data.

E. Martin, Seungwon Lee, K. Riazi et al. · 0 citations
Preprint Jul 2026

MedLLM: An Open Medical Language Model at the Sub-Billion Scale

Across medical benchmarks, MedLLM shows a pattern visible only at sub-billion scale: medical competence does not degrade uniformly under compression but splits by task type and dissociation is masked at 7B, where both capabilities are present, and surfaces only when capacity is scarce.

M. R. Rahman, Asim Ahmed, Mihan Mohagheghzadeh et al. · 0 citations