Skip to content
Open access

Too agreeable to be accurate? Sycophancy and diagnostic instability of large language models in medical diagnosis.

Sep 2026 · Artificial Intelligence in Medicine · Vol 182, pp. 103516 · 0 citations · 18 references
Medicine

TL;DR

The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy.

Abstract

Diagnostic accuracy in Large Language Models (LLM) is an increasing concern as physicians employ LLMs into their medical practice. We measured whether diagnostic correctness in LLMs changed after a certainty challenge or clinician specialty framing. To evaluate the performance of LLMs, we curated 120 public clinical vignettes: 40 MultiCaRe-derived clinical narratives and 80 MedMCQA-derived exam style cases. Ten proprietary and open-weight LLMs were tested across four prompt conditions and three separate runs with two passes per condition, yielding 28,800 responses. Diagnostic correctness was evaluated using an LLM-as-a-judge approach. At neutral baseline, accuracy varied by model and was generally lower for MultiCaRe than MedMCQA. GPT-5 had the highest baseline and post-challenge accuracy (74.4% and 75.3%); Claude Sonnet 4 had the highest accuracy changing flip rate (AcFR; 58.6%) and largest post-challenge accuracy loss (-36.4 percentage points). Across all neutral model-case pairs, baseline accuracy decreased from 51.8% to 42.2% after 'Are you sure?' (AcFR = 32.8%). Correct-to-incorrect transitions (n = 544) exceeded incorrect-to-correct transitions (n = 198). Specialty context prompts produced smaller changes. Adjacent specialty prompts increased overall accuracy by 2.8 percentage points, and differential specialty prompts decreased accuracy by 1.2-1.8 percentage points. The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy. Diagnostic LLM evaluations should report first pass accuracy, harmful flips, beneficial corrections, and stability under clinically plausible conversational challenges before clinical deployment.

Read PDF

Similar papers

Open access Sep 2026

Diagnostic Accuracy and Alignment With Human Reader Responses Across Five Large Language Models on Complex Clinical Case Challenges

Introduction: Large language models (LLMs) match or exceed physician accuracy on benchmark diagnostic tasks. Whether higher accuracy reflects human-like diagnostic behavior, which bears on how safely clinicians can supervise these systems, is rarely assessed. Objectives: To determine whether diagnostic accuracy and al...

R. Bellocco, L. Soraci, Lorenzo Lo Cicero et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Scaling Clinical Judgment to Evaluate Medical AI

Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific ques...

Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur et al. · 0 citations
Open access Sep 2026

Large Language Model Performance on Multistep Clinical Cases: Comparative Study Across Question and Case Levels.

BACKGROUND Most large language models (LLMs) have achieved passing scores on medical licensing examinations. However, most evaluations focus on single-question accuracy, overlooking performance on multistep patient management scenarios, such as making a diagnosis followed by a treatment plan. It is unclear if LLMs can...

Jia-Xue Cha, Yan Zhao, Hui Zong · 0 citations
#artificial intelligence Preprint Sep 2026

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting....

Abinitha Gourabathina, Hao-Ran Zhang, Yue-Xing Hao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.