Skip to content

Author

Şule Asri

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Sep 2026

Phase-Based Dynamic Prediction of Delayed Cerebral Ischemia After Aneurysmal Subarachnoid Hemorrhage: Comparison of a Large Language Model with Intensive Care Specialists

Background and Objectives: Delayed cerebral ischemia (DCI) is a major determinant of poor outcome after aneurysmal subarachnoid hemorrhage (aSAH), yet early risk prediction remains difficult, particularly in sedated or ventilated intensive care patients. We evaluated whether a large language model (LLM) could predict DCI from phase-based dynamic clinical data, compared with intensive care specialists. Materials and Methods: In this single-center, retrospective study, 216 consecutive patients with aSAH were assessed at three predefined phases of accumulating clinical data (day 1; days 1 + 3; days 1 + 3 + 5). For each patient–phase, an LLM (ChatGPT, GPT-5.5 Thinking) and two blinded intensive care specialists predicted DCI risk using only the data available up to that time point. DCI was adjudicated by a blinded three-member panel. Discrimination was assessed by the area under the receiver operating characteristic curve (AUC), with non-inferiority defined a priori as Δ = 0.10. Calibration, decision-curve analysis, and reproducibility were also assessed. Results: Of 216 patients, 60 (27.8%) were DCI-positive and 15 were indeterminate; the primary sample comprised 201 patients. In the prespecified primary analysis, the LLM met the non-inferiority criterion relative to both specialists across all three phases (LLM AUC 0.703–0.747; specialists 0.712–0.764); however, in an equal-granularity sensitivity analysis, non-inferiority remained supported only in Phases 2 and 3 and was not demonstrated in Phase 1. Discrimination increased numerically as data accumulated (Phase 1 vs. 3, p = 0.072), an increase that was attenuated in a landmark-restricted analysis accounting for DCI-onset timing. The LLM showed the lowest false-reassurance rate (3.3–11.7%), reflecting a more cautious threshold rather than better discrimination. Confidence did not reliably indicate accuracy (~25% of high-confidence predictions were wrong); outputs were highly reproducible (Fleiss κ 0.887; intraclass correlation coefficient (ICC) 0.968). Conclusions: An LLM achieved DCI discrimination that was non-inferior to—but not better than—that of experienced specialists; its unreliable confidence scores support clinician-supervised rather than autonomous use.

Mustafa Ay, Tüfek Öztan Dilara, Şule Asri et al. · 0 citations