Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios
Jul 2026· Pediatric surgery international (Print)· Vol 42· 0 citations· 15 references
Medicine
TL;DR
ChatGPT is currently the most robust model for guideline-based management of UPJO, however, the “deceptive confidence” of models like Copilot poses a risk of misinformation and future integration should explore multimodal capabilities.
ObjectiveLarge language models (LLMs) are increasingly used to answer medical questions, but their reliability in guideline-based urological decisions remains uncertain. This study aimed to evaluate the concordance of three widely used LLMs with the European Association of Urology (EAU) 2026 guideline recommendations for the active management of ureteral stones.Materials and MethodsIn this cross-sectional, vignette-based comparative study, the EAU 2026 ureteral-stone treatment algorithm was converted into 40 standardized clinical vignettes (four groups of ten: proximal 10 mm, distal 10 mm). ChatGPT, Gemini, and Claude were queried with the same standardized prompt, which did not name a specific guideline. Responses were scored against predefined EAU-based reference answers using a binary system. Concordance was compared with Cochran’s Q test.ResultsA total of 120 LLM-generated responses were evaluated. Overall concordance was 96.7% (116/120). ChatGPT achieved complete concordance (40/40, 100%), while Gemini and Claude each achieved 95% (38/40); the difference was not significant (Cochran’s Q=2.67, p=0.264). Concordance was complete in all 10 mm scenarios requiring URS prioritization. All four discordances were “incorrect prioritization” in >10 mm stones, presenting shock-wave lithotripsy as co-equal to ureteroscopy; each involved cross-guideline conflation with American Urological Association (AUA) framing. No unsafe recommendation or guideline hallucination was observed under the predefined scoring categories.ConclusionThe evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones. Concordance was complete in non–priority-sensitive
Unknown authors· Bolu Abant Izzet Baysal Univ...· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
Objective
To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures.
Methods
In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ.
Results
The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage).
Conclusions
LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios.
Clinical Relevance
Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.
S. Okur, Ç. Özkalıpçı, Büşra Baykal et al.· Journal of the American Vete...· 0 citations
ABSTRACT
BACKGROUND: Pelvi-ureteric junction obstruction (PUJO) is a functional or anatomical obstruction at the renal pelvis–ureter junction, leading to impaired urine drainage and is the most common cause of antenatally detected hydronephrosis. Several pyeloplasty prediction scoring systems have been developed to quantify disease severity and estimate the need for surgical intervention. This study aimed to apply these scoring systems not only to stratify severity but also to predict surgical outcomes, allowing more informed parental counselling.
OBJECTIVE: To evaluate the predictive performance of three scoring systems; Pyeloplasty Prediction Score (PPS), Onen Hydronephrosis Grading, and the Hydronephrosis Severity Score (HSS), in foretelling postoperative outcomes in children with PUJO, specifically cortical gain and improvement of hydronephrosis.
MATERIALS & METHODS: This descriptive prospective cross-sectional study included 70 children aged 1–13 years who underwent Anderson–Hynes pyeloplasty for PUJO. Patients with secondary causes or elevated eGFR were excluded. At presentation, all children were categorized into risk groups according to PPS, Onen grade, and HSS based on ultrasonography and MAG-3 renogram findings. Postoperative follow-up at 2 months assessed cortical gain and hydronephrosis resolution.
RESULTS:In 70 children undergoing pyeloplasty, PPS showed superior prediction of hydronephrosis resolution, outperforming HSS (AUC difference –0.229, p=0.002) and Onen (–0.044, p=0.169), while HSS was inferior to Onen (p=0.004). For cortex gain, PPS exceeded HSS (AUC difference 0.214, p=0.040) but was similar to Onen (p=0.730).
CONCLUSION: PPS demonstrated superior predictive value for both hydronephrosis improvement and cortical gain, correlating consistently with risk stratification. Onen and HSS did not show proportional predictive performance for these outcomes at 3-month follow-up.
Subhan Ahmed Sajjad, S. Saulat, Jahanzeb Shaikh et al.· Pakistan journal of medicine...· 0 citations
Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.
Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al.· Clinical and Translational O...· 0 citations