Skip to content

Utility of Large Language Models in Pediatric Trauma Triage: An Age-Stratified Analysis of Prompt Engineering.

Aug 2026 · Journal of Pediatric Surgery · pp. 163385 · 0 citations · 40 references
Medicine

TL;DR

It is found that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups and enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups.

Abstract

Background

Pediatric trauma triage is complicated by age-dependent physiologic variability and low case volumes, contributing to persistent undertriage and overtriage despite standardized guidelines. Large language models (LLMs) may offer potential decision-support regardless of pediatric age group.

Methods

We performed a retrospective analysis of 354 EMS-to-hospital pediatric trauma audio recordings from two Level I pediatric trauma centers (2023-2025). Audio recordings were transcribed and evaluated across four LLM prompting strategies (Simple Zero-Shot, L1-biased, L2-biased, Ensemble), three input formats (raw, processed, structured), and three LLMs. Performance was assessed relative to clinician triage decisions and an Injury Severity Score (ISS ≥16) reference standard. Outcomes included accuracy, sensitivity, specificity, undertriage, overtriage, and agreement (Cohen's κ).

Results

Prompt engineering significantly improved LLM performance. Compared with Simple Zero-Shot prompting (71.5% accuracy), criteria-based strategies achieved higher accuracy, with the Ensemble approach demonstrating the most balanced performance (83.9% vs clinicians 78.8%; p=0.026), equivalent undertriage (31.9% vs 34.0%; p=1.0), and lower overtriage (13.7% vs 19.2%; p=0.022). Performance remained stable across age groups (all p>0.10). Prompt strategy exerted the largest effect on accuracy (Δ21.5%), while model selection (Δ0.7%) and input format (Δ2.2%) had minimal impact. Prompt strategies produced distinct tradeoffs between sensitivity and specificity.

Conclusions

Structured prompt engineering enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups. Prompt design, rather than model selection, was the primary determinant of performance. These findings support that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups.

View source

Similar papers

Review Aug 2026

Clinical evaluation of a vision-language model for optimizing triage and clinical workflows in critical care.

OBJECTIVE Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator. MATERIALS AND METHODS We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications. RESULTS The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training. DISCUSSION VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends. CONCLUSION The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.

I. Strechen, P. Krishnan, O. Kilickaya et al. · 0 citations
Open access Aug 2026

Assessing acuity in pediatric emergency department triage: performance of a large language model

Pediatric triage performance varies across emergency departments (ED), contributing to ongoing challenges in pediatric emergency care. There is growing interest in using large language models (LLMs) to support more consistent triage decision-making in children. We evaluated an LLM’s (GPT-5-mini) ability to identify the higher-acuity child from pairs of de-identified clinical notes. Across 228,104 pediatric ED visits, the LLM achieved an overall accuracy of 0.73 (95% CI, 0.73–0.74) in identifying the higher-acuity child, with accuracy improving as acuity differences between visits increased. The LLM was less likely to be correct when the higher-acuity child was older (odds ratio, 0.62, 95% CI, 0.61–0.63) and when the age difference between children was large (0.75, 95% CI, 0.70–0.79). The LLM showed moderate overall accuracy in assessing pediatric acuity and demonstrated a tendency to prioritize younger children, similar to human performance. These findings highlight the need for pediatric-specific LLM evaluation and optimization before clinical use.

Kush Narang, N. Addo, Christopher Y. K. Williams et al. · 0 citations
Aug 2026

Clinical safety of large language model responses to matched patient-language and clinician-language Turkish obstetric and gynecologic triage prompts: a model-blinded paired-scenario study.

OBJECTIVE To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice. METHODS Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator. RESULTS Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028). CONCLUSION No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.

Onur Ada, Uğurcan Dağlı, E. Bilen et al. · 0 citations
Open access Jul 2026

Evaluation of the performance and temporal variability of large language models in patient education regarding pneumothorax: a seven-day analysis.

This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.

Ömer Önal, Suzan Temiz Bekce · 0 citations