Skip to content

Comparison of large language models in the management of pediatric ureteropelvic junction obstruction: a comparative analysis of 125 clinical scenarios

Jul 2026 · Pediatric surgery international (Print) · Vol 42 · 0 citations · 15 references
Medicine

TL;DR

ChatGPT is currently the most robust model for guideline-based management of UPJO, however, the “deceptive confidence” of models like Copilot poses a risk of misinformation and future integration should explore multimodal capabilities.

View source

Similar papers

Open access Aug 2026

Guideline Concordance of Large Language Models in the Management of Ureteral Stones: A Clinical Vignette-Based Comparative Study

ObjectiveLarge language models (LLMs) are increasingly used to answer medical questions, but their reliability in guideline-based urological decisions remains uncertain. This study aimed to evaluate the concordance of three widely used LLMs with the European Association of Urology (EAU) 2026 guideline recommendations for the active management of ureteral stones.Materials and MethodsIn this cross-sectional, vignette-based comparative study, the EAU 2026 ureteral-stone treatment algorithm was converted into 40 standardized clinical vignettes (four groups of ten: proximal 10 mm, distal 10 mm). ChatGPT, Gemini, and Claude were queried with the same standardized prompt, which did not name a specific guideline. Responses were scored against predefined EAU-based reference answers using a binary system. Concordance was compared with Cochran’s Q test.ResultsA total of 120 LLM-generated responses were evaluated. Overall concordance was 96.7% (116/120). ChatGPT achieved complete concordance (40/40, 100%), while Gemini and Claude each achieved 95% (38/40); the difference was not significant (Cochran’s Q=2.67, p=0.264). Concordance was complete in all 10 mm scenarios requiring URS prioritization. All four discordances were “incorrect prioritization” in >10 mm stones, presenting shock-wave lithotripsy as co-equal to ureteroscopy; each involved cross-guideline conflation with American Urological Association (AUA) framing. No unsafe recommendation or guideline hallucination was observed under the predefined scoring categories.ConclusionThe evaluated LLMs showed high concordance with EAU 2026 first-line treatment recommendations for ureteral stones. Concordance was complete in non–priority-sensitive

Unknown authors · 0 citations
Aug 2026

Large language models demonstrate variable diagnostic performance and a systematic risk of undertriage in surgical triage of feline metacarpal and metatarsal fractures.

Objective To evaluate the diagnostic performance of multiple large language models (LLMs) against expert consensus in determining surgical intervention needs for feline metacarpal and metatarsal fractures. Methods In this retrospective study (December 2023 to February 2025), 73 clinical cases of feline metacarpal and metatarsal fractures were evaluated. Two board-certified veterinary orthopedic surgeons established a reference standard for surgical versus conservative management. Five LLMs (ChatGPT, version 5.2 [OpenAI Inc]; Gemini, version 3 Pro [Alphabet Inc]; Grok, version 4.1 [SpaceXAI]; Qwen, version 3.5 [Alibaba Cloud]; and Claude Sonnet, version 4.5 [Anthropic PBC]; and Claude Sonnet, version 4.5 [Anthropic PBC]) assessed anonymized case summaries using a standardized zero-shot prompt. Model recommendations were compared with the reference standard to calculate accuracy, sensitivity, specificity, and the Cohen κ. Results The reference standard classified 49 cases (67.1%) as surgical and 24 (32.9%) as conservative. ChatGPT achieved the highest performance (accuracy, 84.9%; sensitivity, 79.6%; specificity, 95.8%; κ = 0.69). Other models showed lower performance; Qwen and Claude Sonnet failed to identify any surgical cases (0% sensitivity). A systematic bias toward conservative management was observed across all models, causing high false-negative rates (undertriage). Conclusions LLMs demonstrate highly variable diagnostic performance and a systematic risk of undertriage in feline fracture assessment. While top-performing models approach expert-level agreement, others fail in critical clinical scenarios. Clinical Relevance Independent LLM use for surgical decision-making in feline orthopedics is not recommended due to undertriage risks. These tools require strict clinician oversight for preliminary triage.

S. Okur, Ç. Özkalıpçı, Büşra Baykal et al. · 0 citations
Open access Jul 2026

The Scorecard: Comparing different scoring systems to predict the surgical outcome of pyeloplasty in children

ABSTRACT BACKGROUND: Pelvi-ureteric junction obstruction (PUJO) is a functional or anatomical obstruction at the renal pelvis–ureter junction, leading to impaired urine drainage and is the most common cause of antenatally detected hydronephrosis. Several pyeloplasty prediction scoring systems have been developed to quantify disease severity and estimate the need for surgical intervention. This study aimed to apply these scoring systems not only to stratify severity but also to predict surgical outcomes, allowing more informed parental counselling. OBJECTIVE: To evaluate the predictive performance of three scoring systems; Pyeloplasty Prediction Score (PPS), Onen Hydronephrosis Grading, and the Hydronephrosis Severity Score (HSS), in foretelling postoperative outcomes in children with PUJO, specifically cortical gain and improvement of hydronephrosis. MATERIALS & METHODS: This descriptive prospective cross-sectional study included 70 children aged 1–13 years who underwent Anderson–Hynes pyeloplasty for PUJO. Patients with secondary causes or elevated eGFR were excluded. At presentation, all children were categorized into risk groups according to PPS, Onen grade, and HSS based on ultrasonography and MAG-3 renogram findings. Postoperative follow-up at 2 months assessed cortical gain and hydronephrosis resolution. RESULTS:In 70 children undergoing pyeloplasty, PPS showed superior prediction of hydronephrosis resolution, outperforming HSS (AUC difference –0.229, p=0.002) and Onen (–0.044, p=0.169), while HSS was inferior to Onen (p=0.004). For cortex gain, PPS exceeded HSS (AUC difference 0.214, p=0.040) but was similar to Onen (p=0.730). CONCLUSION: PPS demonstrated superior predictive value for both hydronephrosis improvement and cortical gain, correlating consistently with risk stratification. Onen and HSS did not show proportional predictive performance for these outcomes at 3-month follow-up.

Subhan Ahmed Sajjad, S. Saulat, Jahanzeb Shaikh et al. · 0 citations
Review Aug 2026

ChatGPT-4o as a decision-support tool in a urological tumour board: a prospective evaluation.

Final recommendation concordance did not meet the protocol-defined benchmark, ChatGPT-4o never altered an MTB decision, and clinically relevant errors occurred even among highly concordant outputs, showing that concordance alone does not guarantee safety.

Javier De la Torre-Trillo, Albert Munuera, M. D. Ureña et al. · 0 citations