Skip to content
Open access

Large language models for interpretation of health checkup results

Jul 2026 · npj Digital Medicine · Vol 9 · 0 citations · 46 references
Medicine

TL;DR

Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Abstract

Large language models (LLMs) show strong generalization, yet their ability to interpret structured medical data remains insufficiently studied. This work evaluated four LLMs—Claude Sonnet 4, Gemini 2.5 Pro, GPT-4o, and LLaMA 3.1-70B—using comprehensive health checkup data from the Korean National Health Insurance Service. Multiple prompting strategies (few-shot, role-based, constraint-based, and Chain-of-Thought) were tested. Zero-shot accuracy averaged 0.69 (SD 0.06), increasing to 0.92 (0.06) with combined strategies and to 0.95 (0.07) with Chain-of-Thought. Claude Sonnet 4, Gemini 2.5 Pro, and GPT-4o achieved the highest accuracies (≥ 0.98), while LLaMA 3.1-70B showed lower but improvable performance. Item-level analysis of 10,000 cases demonstrated near-perfect accuracy (0.99–1.00) for most biochemical markers, including glucose, cholesterol, triglycerides, and liver enzymes. In contrast, blood pressure showed lower accuracy (0.61–0.91), with age-related decline, likely due to the complexity of multi-categorical thresholds requiring integration of systolic and diastolic values. Subgroup analyses revealed model-specific biases: sex-related biases were observed in body mass index (Claude Sonnet 4) and urine protein, serum creatinine, and gamma-glutamyl transferase (LLaMA 3.1-70B), while age-related biases were identified in blood pressure (Claude Sonnet 4, Gemini 2.5 Pro) and low-density lipoprotein cholesterol and hemoglobin (LLaMA 3.1-70B). Overall, advanced prompting markedly improved model performance, and top-tier LLMs demonstrated robust interpretive capability, while caution is needed for variables with complex clinical semantics such as blood pressure.

Read PDF

Similar papers

Review Open access Jul 2026

A Large Language Model Leaderboard for Clinical Note Entity Extraction

An LLM leaderboard showing how open-source LLMs perform at entity extraction on unseen clinical notes is developed, showing that large language models are already available that can perform entity extraction well enough to be considered in place of some administrative data.

E. Martin, Seungwon Lee, K. Riazi et al. · 0 citations
Open access Jul 2026

The potential of LLMs in generating questions and answers with EHRs

Conventional medical education requires clinicians to formulate questions and answers based on prototypes from EHRs, which is heuristic and time-consuming, this study shows that mainstream LLMs could generate questions and answers with real-world EHRs at levels close to clinicians.

Yunqi Zhu, Wen Tang, Huayu Yang et al. · 1 citation
Open access Aug 2026

The use of large language models in automated depression detection.

Current performance estimates of LLMs with respect to depression screening are most likely optimistic, but when restricted to smaller models that could be locally deployed (for privacy protection) in a clinical setting, LLMs do not detect depression with sufficient accuracy, sensitivity, or specificity to be used in a screening programme.

S. H. Ling, W. Chorney · 0 citations
Review Open access Aug 2026

Explainability of decoder-only clinical large language models: A scoping review.

Findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited, and three regulatory priorities are highlighted: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.

Nishant Mishra, Ameen Abu-Hanna, Iacer Calixto · 0 citations
Review Open access Aug 2026

Robustness Gap of Large Language Models in Nephrology

Background: Whether benchmark performance reflects robust clinical reasoning rather than surface-level pattern recognition remains uncertain. We evaluated the robustness of state-of-the-art large language models (LLMs) on nephrology board renewal questions using "None of the other answers" (NOTA) substitution. Methods: From 210 Japanese Society of Nephrology board renewal questions (2014-2023), two nephrologists independently reviewed all items. Questions in which NOTA became the sole correct answer after replacement were included, yielding 145 validated questions. GPT-5, GPT-4o, Gemini 2.5 Pro, and Gemini 2.0 Flash were evaluated via application programming interfaces under default settings. The primary endpoint was accuracy, and paired differences were assessed using the exact two-sided McNemar test. Results: Accuracy was significantly lower after NOTA substitution for all models: GPT-4o, 66.21% to 19.31% (drop, 46.90 percentage points [pp]); GPT-5, 87.59% to 73.10% (14.48 pp); Gemini 2.0 Flash, 58.62% to 31.03% (27.59 pp); and Gemini 2.5 Pro, 86.90% to 55.86% (31.03 pp); all P < .001. GPT-5 showed the smallest decline and the highest accuracy in both versions. Conclusions: All evaluated LLMs showed a significant robustness gap after NOTA replacement. Newer models may be more robust, but multiple-choice accuracy remains an incomplete measure of clinical reasoning robustness.

A. Soejima, F. Kitano, D. Ichikawa et al. · 0 citations