Laboratory Medicine Decision Support—Beyond Exam Passing: A Blinded 100-Case Text-Based Benchmark of Diagnostic Accuracy, Management Quality, and Safety for ChatGPT, Gemini, and DeepSeek—LLM Decision Support in Laboratory Medicine
ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy.
Abstract
Background/Objective: Large language models (LLMs) have shown exam-level performance, yet their reliability and safety in laboratory medicine—where quantitative data interpretation is central—remain insufficiently validated. This study compared the accuracy, interpretive quality, and safety of ChatGPT-5.2, Gemini 3 Pro, and DeepSeek-V3.2 using a standardized, text-based educational benchmark of clinical pathology and laboratory medicine vignettes. Methods: For each case, the original open-ended questions were answered by each model and scored by two blinded expert raters across four domains—diagnostic accuracy, interpretation, management/investigations, and safety—using a six-point Likert scale, yielding a composite score ranging from 4 to 24. Investigators developed five single-best-answer MCQs per case (500 MCQs total) with consensus answer keys; models selected one option per item. Results: Inter-rater agreement was high (κ = 0.80 for diagnostic concordance; κ = 0.78 for safety flags). Mean composite open-ended scores were 22.5 ± 2 for ChatGPT-5.2, 21.1 ± 2.5 for Gemini, and 20.6 ± 2.6 for DeepSeek (p < 0.001). Fully concordant primary diagnoses were observed in 88%, 84% and 82% of cases, respectively. Unsafe or guideline-discordant recommendations were uncommon but present (2%, 4%, 5%). MCQ accuracy was high: 96% (480/500) for ChatGPT-5.2, 95% (475/500) for Gemini, and 94% (470/500) for DeepSeek. Conclusions: All three LLMs achieved high performance in this standardized text-based benchmark. ChatGPT-5.2 had the highest observed performance across several predefined outcomes, although absolute differences were modest for some measures, particularly MCQ accuracy. These findings should not be interpreted as evidence of clinical superiority or real-world effectiveness.
This study provides a benchmark of LLM performance for radiology OSCE-style content generation and evaluation during a specific snapshot of artificial intelligence development (August 2024).
Ankush Ankush, Samriddhi Burman, Sydney Smith et al.· Radiology Advances· 0 citations
This study aimed to compare the performance of two large language models (LLMs) the ChatGPT-4 and the Google Bard in delivering medical information about colonoscopy. It also sought to evaluate the models’ responses to common patient and caregiver questions in terms of accuracy, applicability, comprehensiveness, and co...
Neslihan Güneş Aydemi̇r, D. Yapar, Yasemin Demi̇r Avcı et al.· BMC Gastroenterology· 0 citations
The hypMOOVE-PHC study is the hypothetical vignette phase of the Massive Open Online Validation and Evaluation (MOOVE) initiative, implemented in Kenya, Malawi, and Tanzania, and aims to validate a pool of LLMs through clinical review of expert-generated vignettes through fully crossed repeated-measures comparative eva...
P. Macharia, C. Kachimanga, M. Mahende et al.· medRxiv· 0 citations
Patients recovering from lumbar fusion increasingly seek guidance from large language model (LLM) chatbots when their surgeon is unavailable, but the accuracy, completeness, and safety-related content of such responses have not been compared across models for the postoperative period.
To compare the accuracy...
Ahmet Kürşat Kara, Ozan Işık· Frontiers in Surgery· 0 citations
ChatGPT Plus achieved overall diagnostic accuracy close to routine radiologist reports, with substantial agreement and high specificity but lower sensitivity, particularly in extremity radiographs.
Uzma Khan, Shravani Moholkar, Fatema Kazi et al.· Indian Journal of Radiology...· 0 citations
Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed.
Objectives: To determine whether diagnostic accuracy and al...
R. Bellocco, L. Soraci, Lorenzo Lo Cicero et al.· Epidemiology Biostatistics a...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.