Jul 2026· International Journal for Research in Applied Science and Engineering Technology· Vol 14, pp. 1921-1929· 0 citations
TL;DR
Findings show that zero-shot language models can provide discrimination comparable to the strongest modified clinical score in this dataset, and mean that the models should be considered complementary research tools, not stand-alone clinical detectors.
Abstract
Timely recognition of sepsis remains difficult when early physiological abnormalities are subtle or incomplete. This
study examined whether general-purpose large language models could discriminate sepsis risk from an initial ICU vital-sign
snapshot as effectively as established clinical scoring approaches. We performed a retrospective benchmark using the deidentified PhysioNet Sepsis Prediction Dataset. The eligible cohort included 39,234 adults, of whom 2,733 were sepsis-positive.
Two zero-shot language models and two modified clinical scores were evaluated on the same patients. Discrimination was
measured by the area under the receiver operating characteristic curve (AUROC), with bootstrap confidence intervals and
DeLong tests for paired comparisons. AUROC was 0.597 for GPT-5.5, 0.592 for modified NEWS2, 0.591 for Claude Sonnet 5,
and 0.570 for modified SOFA. Neither language model differed significantly from modified NEWS2, whereas both produced
higher AUROC values than modified SOFA. A pre-onset-only sensitivity analysis, restricted to patients whose snapshot preceded
their first positive sepsis label (achieved median lead time 39 hours), showed all four estimators converging (AUROC 0.579–
0.591) with no significant pairwise differences, indicating the primary-analysis gap over SOFA was partly attributable to postonset records. A stricter analysis limited to a minimum 6-hour lead time reversed the ranking: the SOFA-derived score obtained
the highest AUROC (0.597), with the LLMs and NEWS2-derived score converging lower (0.573–0.578), again with no
significant pairwise differences. Their alerting behaviour was not interchangeable: GPT-5.5 produced a more balanced
sensitivity-specificity profile, while Claude Sonnet 5 identified more positive cases at the cost of additional false alerts. These
findings show that zero-shot language models can provide discrimination comparable to the strongest modified clinical score in
this dataset. The modest absolute AUROC values, use of modified scores, and retrospective single-dataset design mean that the
models should be considered complementary research tools, not stand-alone clinical detectors.
STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality, outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters.
I. K. Kalyvianakis, C. De Amezaga, E. S. Lee et al.· medRxiv· 0 citations
Sepsis remains a major cause of morbidity and mortality in Intensive Care Units (ICUs). Timely identification of sepsis can prevent severe complications by enabling early treatment, such as administering antibiotics. Despite advances in diagnostic biomarkers and scoring systems, these approaches often lack the ability...
Charithea Stylianides, Andria Nicolaou, Anna Vavlitou et al.· IEEE journal of biomedical a...· 0 citations
Background Pediatric sepsis continues to pose a major challenge in healthcare, compounded by delayed diagnosis and treatment resulting in poor outcomes. Artificial intelligence (AI) and machine learning (ML) continue to develop predictive models that can support the early identification of pediatric sepsis and assist w...
A. Shibu, J. Aadhira, S. Mitra et al.· Journal of Drug Delivery and...· 0 citations
Background Diagnosing infections remains challenging. Clinicians rely on scoring systems and experience, but artificial intelligence (AI) is used to support decision-making by integrating clinical data. However, most AI models focus on predicting adverse outcomes (e.g., ICU admission or sepsis) rather than differentiat...
Sara N. Søgaard, H. Skjøt-Arkil, C. Mogensen et al.· Frontiers in Artificial Inte...· 0 citations
Sepsis remains a major public health concern and is associated with high mortality. Early detection and timely intervention are critical for improving outcomes, yet no standardised approach has been universally adopted. The LiSep LSTM model, a Long Short-Term Memory neural network developed using the MIMIC-III database...
Hong-Yong Tan, L. Froese, D. Wilhelms et al.· Scientific Reports· 0 citations
An interpretable RF-based machine learning model for predicting in-hospital mortality in patients with sepsis was developed and internally validated and showed acceptable discrimination, reasonable calibration, and potential clinical net benefit in internal validation.
Tian-Yu Zhao, Ke-Xin Wen, Xu-Min Han et al.· Frontiers in Medicine· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.