Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.
Abstract
Large language models (LLMs) show strong generalization, yet their ability to interpret structured medical data remains insufficiently studied. This work evaluated four LLMs—Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, and LLaMA 3.1-70B—using comprehensive health checkup data from the Korean National Health Insurance Service. Multiple prompting strategies (few-shot, role-based, constraint-based, and Chain-of-Thought) were tested. Zero-shot accuracy averaged 0.69 (SD 0.06), increasing to 0.92 (0.06) with combined strategies and to 0.95 (0.07) with Chain-of-Thought. Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o achieved the highest accuracies (≥ 0.98), while LLaMA 3.1-70B showed lower but improvable performance. Item-level analysis of 10,000 cases demonstrated near-perfect accuracy (0.99–1.00) for most biochemical markers, including glucose, cholesterol, triglycerides, and liver enzymes. In contrast, blood pressure showed lower accuracy (0.61–0.91), with age-related decline, likely due to the complexity of multi-categorical thresholds requiring integration of systolic and diastolic values. Subgroup analyses revealed model-specific biases: sex-related biases were observed in body mass index (Claude Sonnet 4) and urine protein, serum creatinine, and gamma-glutamyl transferase (LLaMA 3.1-70B), while age-related biases were identified in blood pressure (Claude Sonnet 4, Gemini 2.5 Pro) and low-density lipoprotein cholesterol and hemoglobin (LLaMA 3.1-70B). Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.
An LLM leaderboard showing how open-source LLMs perform at entity extraction on unseen clinical notes is developed, showing that large language models are already available that can perform entity extraction well enough to be considered in place of some administrative data.
E. Martin, Seungwon Lee, K. Riazi et al.· International Journal of Pop...· 0 citations
Although contemporary LLMs increasingly reflect medical consensus for CNS metastases, inconsistent reliability remains a concern, underscoring the need for caution in patient use.
Michael Fiorino, Mei Hainline, Tanay Poddar et al.· Neuro-Oncology Advances· 0 citations
Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.
Yunqi Zhu, Wen Tang, Huayu Yang et al.· Frontiers in Digital Health· 1 citation
Current performance estimates of LLMs with respect to depression screening are most likely optimistic, but when restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.
S. H. Ling, W. Chorney· Acta Psychologica· 0 citations
Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal questions using "None of the other answers" (NOTA) substitution. Methods: From 210 Japanese Society of Nephrology board renewal questions (2014-2023), two nephrologists independently reviewed all items. Questions in which NOTA became the sole correct answer after replacement were included, yielding 145 validated questions. GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash were evaluated via application programming interfaces under default settings. The primary endpoint was accuracy, and paired differences were assessed using the exact two-sided McNemar test. Results: Accuracy was significantly lower after NOTA substitution for all models: GPT-4o, 66.21% to 19.31% (drop, 46.90 percentage points [pp]); GPT-5, 87.59% to 73.10% (14.48 pp); Gemini 2.0 Flash, 58.62% to 31.03% (27.59 pp); and Gemini 2.5 Pro, 86.90% to 55.86% (31.03 pp); all P < .001. GPT-5 showed the smallest decline and the highest accuracy in both versions. Conclusions: All evaluated LLMs showed a significant robustness gap after NOTA replacement. Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness.
A. Soejima, F. Kitano, D. Ichikawa et al.· medRxiv· 0 citations