Skip to content

Diagnostic and Clinical Management Performance of Large Language Models in Pediatric Infectious Diseases: A Multi-Model Comparative Cross-Sectional Study with Healthcare Professionals and Medical Students

Jul 2026 · Journal of Pediatric Infectious Diseases · Vol 21 · 0 citations

TL;DR

Current-generation LLMs demonstrated superior performance in diagnostic accuracy and management score compared with healthcare professionals and medical students on standardized pediatric infectious disease cases, and these findings support their potential role as clinical decision support tools.

Abstract

Objective: Large language models (LLMs) are increasingly investigated to support clinical decision-making. However, their comparative performance against human expertise in pediatric infectious diseases remains insufficiently characterized. This study compared diagnostic accuracy and management score between healthcare professionals, medical students, and multiple LLM configurations. Methods: A cross-sectional comparative study was conducted using 30 standardized pediatric infectious disease cases stratified by difficulty (6 easy, 12 medium, 12 difficult). Each case included three multiple-choice questions: a diagnostic gateway question and two management questions. A conditional scoring framework was applied: management questions were scored only when the diagnostic gateway question was answered correctly. The human cohort included 308 participants, generating 2,450 responses. The LLM cohort included 14 model configurations: ChatGPT (5.1 and 5.2, each in Instant and Thinking modes; 5.3 Instant; 5.4 Thinking); Claude (Opus 3, 4.5, 4.6); Gemini (3.1 Pro, Thinking, Fast); and DeepSeek v3 (Normal, DeepThink). Each was evaluated across five independent sessions (2,100 case responses). Results: The LLM cohort demonstrated significantly higher performance than the human cohort across all endpoints. The adjusted difference in diagnostic accuracy was +25.78 percentage points (95% CI: 22.05–29.52; p < 0.001). The overall score was higher by +36.57 points (95% CI: 32.65–40.50; p < 0.001), and the management score by +22.82 points (95% CI: 19.79–25.85; p < 0.001). Among human participants, pediatric specialists achieved the highest performance (overall score: 79.34); no subgroup reached LLM-level performance. Human performance declined with increasing case difficulty, whereas LLM performance remained relatively stable. Variability among LLM configurations was minimal (range: 98.00%–100.00%). Conclusion: Current-generation LLMs demonstrated superior performance in diagnostic accuracy and management score compared with healthcare professionals and medical students on standardized pediatric infectious disease cases. These findings support their potential role as clinical decision support tools. However, further studies are required to evaluate real-world applicability, safety, and human–AI collaboration in clinical practice.

View source

Similar papers

Open access Jul 2026

Comparing the text-based diagnostic reasoning performance of emergency medicine physicians and large language models in both definitive and differential diagnoses using standardized clinical vignettes: a preliminary study

Large language models outperformed emergency medicine physicians in overall diagnostic accuracy and demonstrated superior consistency across different types of diagnostic tasks, indicating that AI possesses robust pattern-recognition and reasoning capabilities, suggesting it could serve as a highly reliable clinical decision support tool.

Mehdi Arzani Shamsabadi, Roya Vatankhah, Hasan Jalilvand et al. · 0 citations
Open access Jul 2026

Diagnostic capability of large language models in critically ill patients: a prospective single-centre study comparing ChatGPT, Claude, and Gemini with emergency physicians.

BACKGROUND Clinical decision-making requires integrating history, physical examination, laboratory, and imaging data. In the emergency department (ED), workload, time pressure, and cognitive burden may impair this process and affect decision quality. This study compares the diagnostic outputs of ChatGPT, Claude, and Gemini with those of emergency physicians in real-world ED cases. METHODS This prospective, single-centre observational diagnostic agreement study compared the stage-wise outputs of four Large Language Models (LLMs) (ChatGPT-4o, ChatGPT-5, Claude Opus 4.1, and Gemini 2.5 Pro) with those of emergency physicians in critically ill ED patients. Between 10 August and 10 September 2025, de-identified clinical data were entered into the models via their official web interfaces using standardised prompts. In the first stage, physicians and LLMs each generated five preliminary diagnoses based on vital signs and medical history. In the second stage, following physical examination and laboratory and imaging results, both refined their lists into three differential diagnoses. In the third stage, the physicians' final diagnosis was accepted as the reference, and each LLM was prompted to provide a final diagnosis. LLM preliminary and differential diagnoses were compared with those of the physicians at the corresponding stage, and LLM final diagnoses with the reference; the inclusion of the final diagnosis within earlier lists was also evaluated. Agreement was quantified using Cohen's κ; analyses were performed in R. RESULTS Of 389 screened patients, 180 were included (56.1% male; mean age 67 ± 15.9 years). Physicians contained the reference diagnosis within their top-5 preliminary and top-3 differential lists in 83.9% and 98.3% of cases, respectively, significantly exceeding every LLM (all p < 0.001). Final-diagnosis match rates were 67.2% [60.3-73.5] for ChatGPT-4o, 65.6% [58.7-71.9] for ChatGPT-5, 63.3% [56.3-69.9] for Claude Opus 4.1, and 59.4% [52.3-66.1] for Gemini 2.5 Pro (p = 0.16). Cohen's κ ranged from 0.575 (Gemini 2.5 Pro) to 0.656 (ChatGPT-4o), indicating moderate-to-substantial agreement, with no pairwise difference reaching significance. CONCLUSIONS The LLMs achieved moderate agreement with ED reference diagnoses in critically ill patients but were consistently outperformed by physicians at the early diagnostic phases. Despite final-diagnosis match rates of 59%-67%, their current diagnostic role in the ED remains limited.

İbrahim Günaydın, M. Yılmaz, Sinan Akpunar et al. · 0 citations
Open access Aug 2026

Stepwise Diagnostic Evaluation of Chinese Large Language Models: Comparative Study of Common and Rare Diseases

Abstract Background Large language models (LLMs) are increasingly applied in clinical decision support, yet their diagnostic performance in Chinese-language settings and under realistic clinical workflows remains unclear. In particular, how LLMs perform across diseases with different prevalence and under stepwise diagnostic processes has not been well characterized. Objective This study aimed to evaluate the diagnostic capabilities of LLMs for common diseases and rare diseases using clinical vignettes within a hypothetico-deductive framework and to identify their potential and limitations for clinical diagnosis. Methods We evaluated 4 Chinese LLMs (Doubao 1.5, DeepSeek-V3, Kimi K1.5, and Leftdoctor GPT 3.5) using 56 clinical cases (28 chronic obstructive pulmonary disease [COPD], and 28 relapsing polychondritis [RP]) sourced from the China Clinical Case Results Database (March 31-April 14, 2025). Patient information was provided incrementally, starting with the initial medical history, followed by physical examination, and laboratory results. Evaluation metrics included top-3 accuracy (RTop3D), top-1 accuracy (RTopD), final diagnostic accuracy (RFA), and mean reciprocal rank (MRR). Statistical analysis was performed using generalized estimating equations (GEE), Friedman tests, and Wilcoxon signed-rank tests with Bonferroni correction. In addition, a qualitative analysis was conducted to characterize recurrent patterns of diagnostic errors. Results LLMs demonstrated significantly higher diagnostic accuracy for COPD compared to RP across all metrics (P<.001). Diagnostic accuracy improved after additional clinical information was provided, with the improvement mainly observed in RP cases. In RP, diagnostic accuracy increased from 32.14% to 71.43% for DeepSeek and from 35.71% to 78.57% for Doubao, whereas COPD accuracy remained consistently high across all diagnostic stages (82.14%‐92.86%). For COPD, ranking performance was high and comparable among all models (MRR range: 0.82‐0.89; P=.71). In RP, diagnostic performance differed significantly among models (MRR range: 0.10‐0.39; P<.001). Qualitative analysis showed that COPD errors were mainly related to a failure to recognize specific features, whereas RP errors involved more diverse patterns, particularly the neglect of negative evidence and the failure to recognize specific features. Conclusions Chinese LLMs demonstrated relatively strong diagnostic performance for common diseases such as COPD, but lower and less stable performance for rare diseases such as RP. Additional clinical information improved diagnostic accuracy primarily in RP cases, although differences between models remained evident under diagnostically complex conditions. Error patterns in RP cases suggest that current LLMs remain limited in their ability to integrate complex clinical information and exclusionary findings. Careful evaluation and appropriate clinical oversight remain important for their application in clinical practice.

Jiayi Wang, Jiao Yang, Rui Guo · 0 citations
Open access Jul 2026

Evaluation of four large language models on complex, infectious disease case scenarios

On complex ID scenarios, large language models responses were variable and caution is required when deploying these models in ID domains without specialist oversight, suggesting caution is required when deploying these models in ID domains without specialist oversight.

A. Pradhan, B. Waxse, W. Matias et al. · 0 citations
Jul 2026

The illusion of competence: Evaluating the clinical reasoning of large language models in pediatric gastroenterology.

While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.

Y. Ergen, S. Teke, E. G. Başaran et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations