Sep 2026· Artificial Intelligence in Medicine· Vol 182, pp.
103516
· 0 citations· 18 references
Medicine
TL;DR
The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy.
Abstract
Diagnostic accuracy in Large Language Models (LLM) is an increasing concern as physicians employ LLMs into their medical practice. We measured whether diagnostic correctness in LLMs changed after a certainty challenge or clinician specialty framing. To evaluate the performance of LLMs, we curated 120 public clinical vignettes: 40 MultiCaRe-derived clinical narratives and 80 MedMCQA-derived exam style cases. Ten proprietary and open-weight LLMs were tested across four prompt conditions and three separate runs with two passes per condition, yielding 28,800 responses. Diagnostic correctness was evaluated using an LLM-as-a-judge approach. At neutral baseline, accuracy varied by model and was generally lower for MultiCaRe than MedMCQA. GPT-5 had the highest baseline and post-challenge accuracy (74.4% and 75.3%); Claude Sonnet 4 had the highest accuracy changing flip rate (AcFR; 58.6%) and largest post-challenge accuracy loss (-36.4 percentage points). Across all neutral model-case pairs, baseline accuracy decreased from 51.8% to 42.2% after 'Are you sure?' (AcFR = 32.8%). Correct-to-incorrect transitions (n = 544) exceeded incorrect-to-correct transitions (n = 198). Specialty context prompts produced smaller changes. Adjacent specialty prompts increased overall accuracy by 2.8 percentage points, and differential specialty prompts decreased accuracy by 1.2-1.8 percentage points. The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy. Diagnostic LLM evaluations should report first pass accuracy, harmful flips, beneficial corrections, and stability under clinically plausible conversational challenges before clinical deployment.
Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed.
Objectives: To determine whether diagnostic accuracy and al...
R. Bellocco, L. Soraci, Lorenzo Lo Cicero et al.· Epidemiology Biostatistics a...· 0 citations
A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.
Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al.· Indian Journal of Computer S...· 0 citations
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al.· 0 citations
It is suggested that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.
BACKGROUND
Most large language models (LLMs) have achieved passing scores on medical licensing examinations. However, most evaluations focus on single-question accuracy, overlooking performance on multistep patient management scenarios, such as making a diagnosis followed by a treatment plan. It is unclear if LLMs can...
Jia-Xue Cha, Yan Zhao, Hui Zong· JMIR Medical Education· 0 citations
As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting....
Abinitha Gourabathina, Hao-Ran Zhang, Yue-Xing Hao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.