Skip to content

Author

A. Soejima

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Review Open access Aug 2026

Robustness Gap of Large Language Models in Nephrology

Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal questions using "None of the other answers" (NOTA) substitution. Methods: From 210 Japanese Society of Nephrology board renewal questions (2014-2023), two nephrologists independently reviewed all items. Questions in which NOTA became the sole correct answer after replacement were included, yielding 145 validated questions. GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash were evaluated via application programming interfaces under default settings. The primary endpoint was accuracy, and paired differences were assessed using the exact two-sided McNemar test. Results: Accuracy was significantly lower after NOTA substitution for all models: GPT-4o, 66.21% to 19.31% (drop, 46.90 percentage points [pp]); GPT-5, 87.59% to 73.10% (14.48 pp); Gemini 2.0 Flash, 58.62% to 31.03% (27.59 pp); and Gemini 2.5 Pro, 86.90% to 55.86% (31.03 pp); all P < .001. GPT-5 showed the smallest decline and the highest accuracy in both versions. Conclusions: All evaluated LLMs showed a significant robustness gap after NOTA replacement. Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness.

A. Soejima, F. Kitano, D. Ichikawa et al. · 0 citations