Aug 2026· Journal of Pediatric Surgery· pp.
163385
· 0 citations· 40 references
Medicine
TL;DR
It is found that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups and enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups.
Abstract
Background
Pediatric trauma triage is complicated by age-dependent physiologic variability and low case volumes, contributing to persistent undertriage and overtriage despite standardized guidelines. Large language models (LLMs) may offer potential decision-support regardless of pediatric age group.
Methods
We performed a retrospective analysis of 354 EMS-to-hospital pediatric trauma audio recordings from two Level I pediatric trauma centers (2023-2025). Audio recordings were transcribed and evaluated across four LLM prompting strategies (Simple Zero-Shot, L1-biased, L2-biased, Ensemble), three input formats (raw, processed, structured), and three LLMs. Performance was assessed relative to clinician triage decisions and an Injury Severity Score (ISS ≥16) reference standard. Outcomes included accuracy, sensitivity, specificity, undertriage, overtriage, and agreement (Cohen's κ).
Results
Prompt engineering significantly improved LLM performance. Compared with Simple Zero-Shot prompting (71.5% accuracy), criteria-based strategies achieved higher accuracy, with the Ensemble approach demonstrating the most balanced performance (83.9% vs clinicians 78.8%; p=0.026), equivalent undertriage (31.9% vs 34.0%; p=1.0), and lower overtriage (13.7% vs 19.2%; p=0.022). Performance remained stable across age groups (all p>0.10). Prompt strategy exerted the largest effect on accuracy (Δ21.5%), while model selection (Δ0.7%) and input format (Δ2.2%) had minimal impact. Prompt strategies produced distinct tradeoffs between sensitivity and specificity.
Conclusions
Structured prompt engineering enables general-domain LLMs to achieve clinician-comparable performance in pediatric trauma triage across age groups. Prompt design, rather than model selection, was the primary determinant of performance. These findings support that structured prompt engineering improves off-the-shelf LLM performance over unstructured prompting across age groups.
OBJECTIVE
Clinical deterioration in hospitalized patients is often preceded by subtle, dynamic physiological changes that are difficult to detect using intermittently charted electronic health record (EHR) data. Our objective was to evaluate the reliability, interpretability, and clinical relevance of a Vision‑Language Model (VLM)-based triage framework that analyzes physiological trend images, by comparing VLM-generated outputs with attending physician assessments as the expert clinical comparator.
MATERIALS AND METHODS
We conducted a single-center expert agreement pilot study including 100 adult patients with two hours of dynamic monitoring data across four vital signs (SpO2, RR, HR, BP). A structured prompt was developed using the Gemini 2.5 Flash model. Two independent reviewers assessed VLM outputs for clinical interpretation, artifact detection, and triage classification. We evaluated reviewer agreement using percent agreement and weighted Cohen's κ. A secondary risk-oriented analysis measured classification concordance and over-triaged cases as lower-risk classifications, while cases underestimating patient acuity were designated as higher-risk misclassifications.
RESULTS
The VLM demonstrated moderate to substantial agreement in triage classification with the attending physician and a low rate of under-triage. The VLM assigned the same patient acuity category as attending physician in 75 % of cases and underestimated acuity in 7 % of cases, compared with 14 % underestimation by the physician in training.
DISCUSSION
VLMs extend generative artificial intelligence capabilities by enabling image‑grounded clinical reasoning and offer signal‑processing capabilities for interpreting time‑stamped physiological trends.
CONCLUSION
The VLM demonstrated reliable clinical interpretation and an acceptable safety profile, however its integration into clinical workflows for early recognition of physiological deterioration and patient acuity assessment requires further rigorous evaluation and comparison to currently used track-and-trigger systems and patient monitoring methods.
I. Strechen, P. Krishnan, O. Kilickaya et al.· International Journal of Med...· 0 citations
Pediatric triage performance varies across emergency departments (ED), contributing to ongoing challenges in pediatric emergency care. There is growing interest in using large language models (LLMs) to support more consistent triage decision-making in children. We evaluated an LLM’s (GPT-5-mini) ability to identify the higher-acuity child from pairs of de-identified clinical notes. Across 228,104 pediatric ED visits, the LLM achieved an overall accuracy of 0.73 (95% CI, 0.73–0.74) in identifying the higher-acuity child, with accuracy improving as acuity differences between visits increased. The LLM was less likely to be correct when the higher-acuity child was older (odds ratio, 0.62, 95% CI, 0.61–0.63) and when the age difference between children was large (0.75, 95% CI, 0.70–0.79). The LLM showed moderate overall accuracy in assessing pediatric acuity and demonstrated a tendency to prioritize younger children, similar to human performance. These findings highlight the need for pediatric-specific LLM evaluation and optimization before clinical use.
Kush Narang, N. Addo, Christopher Y. K. Williams et al.· npj Health Systems· 0 citations
OBJECTIVE
To determine whether presenting matched obstetric and gynecologic triage scenarios as patient-language prompts rather than clinician-language prompts affects expert-rated clinical confidence and the clinical safety of LLM-generated advice.
METHODS
Thirty obstetric and gynecologic scenarios were presented in matched clinician-language and patient-language Turkish formats to four LLMs. Five specialists independently evaluated 240 responses, generating 1,200 ratings. The primary outcome was the Global Clinical Confidence Score (GCCS; 0-2); five secondary outcomes were rated on 1-5 scales. Associations were examined using ordinal logistic generalized estimating equations adjusted for model and evaluator.
RESULTS
Clinically reliable responses (GCCS = 2) accounted for 89.3% of clinician-language and 91.3% of patient-language ratings. Patient-language phrasing was not significantly associated with overall GCCS (cumulative odds ratio 0.78, 95% confidence interval 0.57-1.06; p = 0.115), and the language-by-model interaction was not significant (p = 0.422). Patient-language prompts were associated with fewer GCCS = 0 ratings in a binary generalized estimating equations analysis (odds ratio 0.66, 95% confidence interval 0.46-0.96; p = 0.031), although the exact paired McNemar test was not significant (p = 0.096). After false-discovery-rate correction, patient-language prompts had higher evaluator-level triage appropriateness and clinical applicability scores (both adjusted p = 0.028).
CONCLUSION
No significant difference in overall expert-rated clinical confidence was detected between patient-language and clinician-language prompts.
Onur Ada, Uğurcan Dağlı, E. Bilen et al.· European Journal of Obstetri...· 0 citations
This exploratory scenario-based study demonstrates model-specific, rather than universal, language effects on LLM performance in dental trauma management.
This study investigates the readability, clinical reliability, and temporal consistency of artificial intelligence (AI) chatbots regarding pneumothorax information. A question bank comprising 40 patient-centered queries was deployed across three large language models (ChatGPT, Gemini, Copilot), stratified by two access tiers and two prompting strategies (zero-shot versus the optimized PROMPORT strategy). Queries were replicated longitudinally on Days 1, 3, and 7 under strict session-control protocols. Text accessibility was quantified using five automated readability indices, while two independent, blinded thoracic surgeons evaluated clinical quality using modified DISCERN (mDISCERN), JAMA benchmarks, and PEMAT-P indices. Readability metrics demonstrated absolute structural stability across the tracking intervals (p > 0.05). Unprompted configurations consistently generated complex, high-school-level outputs, whereas the PROMPORT strategy successfully compressed linguistic variances and neutralized chronological algorithmic drift (p > 0.05). Conversely, unprompted architectures exhibited significant temporal volatility in mDISCERN and JAMA profiles (p < 0.05), which was successfully stabilized by optimized prompt constraints. Inter-rater reliability was high across all structural evaluations. In conclusion, while unprompted models exhibit marked baseline linguistic and quality variations, the strategic integration of robust prompt engineering successfully enforces the temporal stability and clarity required for reliable digital public health communication.
Ömer Önal, Suzan Temiz Bekce· Scientific Reports· 0 citations