Skip to content
Open access

Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine

Sep 2026 · Diagnostics · Vol 16 · 0 citations · 24 references
Medicine

TL;DR

ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.

Abstract

Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine—where quantitative data interpretation is central—remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro, and DeepSeek-V3.2 using a standardized, text-based educational benchmark of clinical pathology and laboratory medicine vignettes. Methods: For each case, the original open-ended questions were answered by each model and scored by two blinded expert raters across four domains—diagnostic accuracy, interpretation, management/investigations, and safety—using a six-point Likert scale, yielding a composite score ranging from 4 to 24. Investigators developed five single-best-answer MCQs per case (500 MCQs total) with consensus answer keys; models selected one option per item. Results: Inter-rater agreement was high (κ = 0.80 for diagnostic concordance; κ = 0.78 for safety flags). Mean composite open-ended scores were 22.5 ± 2 for ChatGPT-5.2, 21.1 ± 2.5 for Gemini, and 20.6 ± 2.6 for DeepSeek (p < 0.001). Fully concordant primary diagnoses were observed in 88%, 84% and 82% of cases, respectively. Unsafe or guideline-discordant recommendations were uncommon but present (2%, 4%, 5%). MCQ accuracy was high: 96% (480/500) for ChatGPT-5.2, 95% (475/500) for Gemini, and 94% (470/500) for DeepSeek. Conclusions: All three LLMs achieved high performance in this standardized text-based benchmark. ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy. These findings should not be interpreted as evidence of clinical superiority or real-world effectiveness.

Read PDF

Similar papers

Review Open access Sep 2026

Benchmarking Large Language Model Performance in Generating and Assessing Radiology Objective Structured Clinical Examination

This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).

Ankush Ankush, Samriddhi Burman, Sydney Smith et al. · 0 citations
Open access Oct 2026

Large language models’ performance in answering common patient questions about colonoscopy: an expert-based evaluation

This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and co...

Neslihan Güneş Aydemi̇r, D. Yapar, Yasemin Demi̇r Avcı et al. · 0 citations
Review Open access Sep 2026

Evaluating large language models as clinical decision support tools in primary healthcare settings: Protocol for a multi-country comparative validation study on expert-adjudicated hypothetical vignettes (hypMOOVE-PHC)

The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...

P. Macharia, C. Kachimanga, M. Mahende et al. · 0 citations
#small language model Open access Oct 2026

How well do large language models answer postoperative lumbar fusion questions? A blinded comparative analysis of accuracy, quality, readability, and safety-related content

Patients recovering from lumbar fusion increasingly seek guidance from large language model (LLM) chatbots when their surgeon is unavailable, but the accuracy, completeness, and safety-related content of such responses have not been compared across models for the postoperative period. To compare the accuracy...

Ahmet Kürşat Kara, Ozan Işık · 0 citations
Open access Aug 2026

Diagnostic Accuracy of ChatGPT Plus (GPT-4o) for the Interpretation of Chest and Extremity Radiographs Against Routine Radiologist Reporting: A Single-Centre, Retrospective, Cross-Sectional Study

ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs.

Uzma Khan, Shravani Moholkar, Fatema Kazi et al. · 0 citations
Open access Sep 2026

Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges

Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed. Objectives: To determine whether diagnostic accuracy and al...

R. Bellocco, L. Soraci, Lorenzo Lo Cicero et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.