Skip to content
Open access

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

Aug 2026 · Indian Journal of Computer Science and Technology · 0 citations

TL;DR

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Abstract

Large language models show potential for clinical diagnostic support, but their diagnostic accuracy across diverse real-world patient presentations remains uncertain. We evaluated diagnostic retrieval and ranking using multi-system emergency-department narratives from MIMIC-IV-Ext version 1.0.2, a deidentified research dataset derived from MIMIC-IV and curated for research involving referral, triage and diagnostic prediction. The dataset was selected because it links early clinical information, including presenting complaints, history and initial vital signs, with documented diagnoses derived from routine care. A locked cohort of 995 diagnosis-free vignettes was used, with one protected index primary diagnosis per case. GPT-5.6 Thinking, Claude Sonnet 5, Llama-3.1-8B-Instruct and Mistral-7B-Instruct-v0.3 independently generated exactly three ranked differential diagnoses for every vignette. The principal outcome was concept-equivalent Top-3 accuracy; Top-1 accuracy, mean reciprocal rank, omission rate, strict text matching, system-wise performance and paired statistical comparisons were secondary outcomes. GPT-5.6 achieved the highest Top-1 accuracy (41.7%), Top-3 accuracy (61.1%) and mean reciprocal rank (0.504). Claude ranked second at 39.8%, 55.3% and 0.467, respectively. Llama reached 25.5% Top-1 and 41.3% Top-3 accuracy, while Mistral reached 22.8% and 38.7%. Overall Top-3 outcomes differed significantly across models (Cochran Q=271.73, df=3, p=1.31×10⁻⁵⁸). Under identical clinical inputs and scoring rules, the proprietary models retrieved the documented index diagnosis more often and ranked it higher than the two open-weight models. These findings provide a reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations and establish a baseline for further clinical validation.

Read PDF

Similar papers

Open access Aug 2026

Incremental Diagnostic Value of Clinical Information for Large Language Models Across Multiple Organs: Retrospective Study

Abstract Background Although large language models (LLMs) have demonstrated the ability to generate the impression section from radiology findings automatically, the incremental diagnostic value of clinical information for these models remains unclear. Objective This study aimed to evaluate the incremental diagnostic value of clinical information for LLMs and compare their performance with that of radiologists. Methods This retrospective study included radiology reports from patients with histopathologically confirmed liver, lung, and breast diseases from 2 institutions between October 2021 and February 2025. We defined three progressive information input scenarios: (1) basic patient information and imaging findings, (2) scenario A plus chief complaint or clinical history, and (3) scenario B plus key laboratory results. Scenario-based data were input into 3 general-purpose LLMs (DeepSeek-R1, Gemini 2.5 Pro, and GPT-4o), generating 2709 entries. Diagnostic accuracy was assessed for both benign-malignant differentiation and disease diagnosis, with histopathology serving as the reference standard. Accuracy was compared among scenarios and against radiologist performance using the McNemar test, and P values were adjusted using the Holm-Bonferroni correction for multiple comparisons. Results A total of 301 patients with pathologically confirmed diseases were included (mean age 53.5, SD 12.0 years; women: n=208, 69.1%). In the liver cohort, a numerical trend toward higher accuracy was observed in scenario C compared with scenario A across all 3 models (scenario C range: 72.3%‐76.2% vs scenario A range: 64.4%‐68.3%); these differences did not reach statistical significance after Holm-Bonferroni correction (all adjusted P>.99). Notably, the DeepSeek-R1 model in scenario C achieved the highest diagnostic accuracy (77/101, 76.2%), with no evidence of a difference compared with radiologists (82/101, 81.2%; adjusted P>.99). In contrast, results in the lung and breast cohorts were more heterogeneous. In the lung cohort, GPT-4o achieved its highest accuracy in scenario A for disease diagnosis (68/92, 73.9%), which exceeded its performance in scenario B (64/92, 69.6%) and scenario C (66/92, 71.7%), suggesting that additional clinical information did not confer a consistent benefit. Gemini 2.5 Pro in scenario B achieved the highest accuracy in this cohort (72/92, 78.3%); however, no statistically significant difference was found compared with radiologists (80/92, 87.0%; adjusted P=.25). In the breast cohort, DeepSeek-R1 achieved the numerically highest diagnostic accuracy in scenario A, and it decreased numerically with the addition of laboratory tests, although no significant difference was found between scenarios A and C (73/108, 67.6% vs 71/108, 65.7%; adjusted P>.99). Conclusions While the addition of clinical information was associated with a numeric trend toward higher diagnostic accuracy overall, this trend was heterogeneous across models and disease types, and no statistically significant improvement was demonstrated after adjustment for multiple comparisons.

Jin-Qi Zhang, Xiao-Yi Wang, Yanfeng Zhao et al. · 0 citations
Review Open access Aug 2026

Large Language Models for Differential Diagnosis: A Survey of Performance, Collaboration, and Technical Strategies

Errors in differential diagnosis often arise while clinicians are generating and comparing candidate explanations. This review examines the use of large language models (LLMs) for this part of diagnostic reasoning. Internal medicine and pediatrics are the main focus; evidence from radiology, surgical subspecialties, infectious disease, and mental health is used to examine how findings change across specialties. Reported performance depends on the clinical setting, the quality of the input, the prompt, model adaptation, and the evaluation design. Some studies place LLMs near trainees and find that they produce wider, better-organized differentials. Experienced clinicians, however, remain more reliable overall. Domain adaptation, external knowledge, and interactive workflows have improved performance in specific evaluations, but hallucinations and automation bias remain, alongside unresolved questions of governance. Current evidence therefore supports clinician-supervised use of artificial intelligence (AI) systems rather than autonomous diagnosis, pending prospective and specialty-specific evaluation.

Yun-Jia Wu, Qi Yan, Dingcheng Tian · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Open access Aug 2026

A Human-in-the-Loop Large Language Model System Based on the Model Context Protocol for Differential Diagnosis from Electronic Medical Records and Literature

DDx-Finder is presented, an open-source framework that leverages Model Context Protocol (MCP) servers for direct EMR and literature access, enabling prompt-driven clinical state extraction and reliable case-report re- trieval via generating searching query by LLM, while addressing limitations related to resource demands and privacy concerns.

H. Lim, H. Yi, J. Y. Yoon et al. · 0 citations
Open access Jul 2026

Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study

Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool.

Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al. · 0 citations