Jul 2026· Journal of Pediatric Gastroenterology and Nutrition - JPGN· 0 citations· 19 references
Medicine
TL;DR
While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence.
Abstract
Objective
To evaluate the diagnostic accuracy and clinical reasoning of three frontier large language models (LLMs) across standardized pediatric gastroenterology, hepatology, and nutrition (PGHN) clinical vignettes.
Methods
In this cross-sectional study, 25 fictional PGHN vignettes were developed by one board-certified pediatric gastroenterologist and evaluated using three LLMs: Gemini 3.1 Pro, ChatGPT 5.4 Thinking, and Claude Sonnet 4.6 Extended. A conditional two-step prompting protocol was applied. Three blinded, PGHN-certified co-authors independently scored responses using a structured instrument covering four domains (diagnostic accuracy, management, patient safety, reference quality; 0-2 each) and a global quality of clinical reasoning (QCR) score (1-5 Likert scale). Interobserver reliability was assessed using the intraclass correlation coefficient (ICC). Group differences were analyzed using the Kruskal-Wallis H test with Dunn post-hoc correction.
Results
Interobserver reliability was excellent (ICC: 0.98 for domains; 0.87 for QCR). All three models achieved perfect diagnostic accuracy (median 2.00, interquartile range: 2.00-2.00). Statistically significant intermodel differences were identified in reference quality (p = 0.048) and QCR (p = 0.049). Claude achieved significantly higher QCR scores than ChatGPT (p = 0.043) and demonstrated the highest overall reference quality. Qualitative analysis revealed critical pharmacological dosing errors, contextual blindness, temporal obsolescence, and a high frequency of fabricated citations.
Conclusion
While current LLMs demonstrate good diagnostic pattern recognition in PGHN, reproducible and potentially life-threatening failures in pharmacological reasoning and reference accuracy create a dangerous illusion of competence. These findings suggest that LLMs may support differential diagnosis brainstorming in PGHN but do not establish their safety in real-world clinical scenarios.
Psychological assessment is essential for bariatric surgery candidacy but remains inconsistent and time-consuming. This study compared three large language models (LLMs) in generating structured psychological screening checklists to support standardized preoperative evaluation.
Models A, B, and C generated checklists for five standardized bariatric vignettes. Three blinded psychologists independently rated the outputs on 5-point Likert scales for clinical relevance, completeness, specificity, and usability. Statistical analysis included Jaccard similarity to quantify content overlap and analysis of variance with Tukey
post hoc
tests to compare expert ratings.
Model A produced the most extensive checklists (mean = 35.6 items, standard deviation = 4.2), while Models B and C were more concise (28.4 and 26.7 items, respectively). Model A achieved the highest completeness scores, whereas Model B was rated highest for specificity and usability. Interrater agreement was good to excellent (intraclass correlation coefficient = 0.79–0.87). Moderate content overlap (Jaccard similarity = 0.45–0.63) suggested complementary model strengths. Between-model differences were significant for completeness,
F
(2, 6) = 11.5,
p
= 0.008, and specificity,
F
(2, 6) = 9.8,
p
= 0.013.
LLMs differ significantly in checklist breadth and focus. Strategic model selection can enhance preoperative workflows by balancing comprehensiveness with targeted, actionable assessment. However, deployment should remain clinician-supervised because outputs may vary with model updates, prompt phrasing, and local workflow constraints.
Y. K. Çalışkan, Fatih Başak· Bariatric Surgical Practice...· 0 citations
ObjectiveThis study evaluates the performance of large language models (LLMs)-ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3-in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models.MethodA total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations.Resultso3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30-32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models.ConclusionAI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.
Bünyamin Arı· Health Informatics Journal· 0 citations
This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.
Ömer Önal, Suzan Temiz Bekce· Scientific Reports· 0 citations
OBJECTIVE
Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator.
MATERIALS AND METHODS
We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications.
RESULTS
The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training.
DISCUSSION
VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends.
CONCLUSION
The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.
I. Strechen, P. Krishnan, O. Kilickaya et al.· International Journal of Med...· 0 citations
Large Language Models (LLMs) such as ChatGPT and Gemini are increasingly used to answer medical and psychological questions, yet systematic evaluations across domains with differing reasoning demands remain limited. We present a multi-domain expert evaluation of two state-of-the-art models, ChatGPT Pro (v5.2) and Gemini 3 Pro, across three healthcare domains: Gynecology, Pathology, and Psychology. We curated 300 open-ended, realistic questions, 100 per domain, designed to elicit clinical reasoning, mechanistic interpretation, and conceptual explanation. Responses were independently scored by domain experts using a standardized five-point rubric. Results reveal domain-dependent performance patterns, with both models performing well in guideline-aligned advisory tasks in Gynecology, lower in mechanistic diagnostic contexts in Pathology, and showing divergent strengths in conceptual psychology questions. Quantitative and qualitative analyses highlight recurring limitations in contextual nuance, mechanistic depth, and safety framing. These findings underscore the importance of domainstratified evaluation and expert oversight when deploying LLMs for healthcare information and provide a reproducible framework for future assessments.
Pragna Prahallad, Pranathi Prahallad, Dhrithi P. Desai· 2026 6th International Confe...· 0 citations
MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time, and represents a step towards building safe LLM systems to enhance patient education in psychiatry.
Alexander J. Hish, A. Nagendran, S. Compton· 0 citations