Skip to content
Open access

Evaluating Large Language Models Against Clinical Assessment Frameworks for Early Sepsis Detection in the ICU

Jul 2026 · International Journal for Research in Applied Science and Engineering Technology · Vol 14, pp. 1921-1929 · 0 citations

TL;DR

Findings show that zero-shot language models can provide discrimination comparable to the strongest modified clinical score in this dataset, and mean that the models should be considered complementary research tools, not stand-alone clinical detectors.

Abstract

Timely recognition of sepsis remains difficult when early physiological abnormalities are subtle or incomplete. This study examined whether general-purpose large language models could discriminate sepsis risk from an initial ICU vital-sign snapshot as effectively as established clinical scoring approaches. We performed a retrospective benchmark using the deidentified PhysioNet Sepsis Prediction Dataset. The eligible cohort included 39,234 adults, of whom 2,733 were sepsis-positive. Two zero-shot language models and two modified clinical scores were evaluated on the same patients. Discrimination was measured by the area under the receiver operating characteristic curve (AUROC), with bootstrap confidence intervals and DeLong tests for paired comparisons. AUROC was 0.597 for GPT-5.5, 0.592 for modified NEWS2, 0.591 for Claude Sonnet 5, and 0.570 for modified SOFA. Neither language model differed significantly from modified NEWS2, whereas both produced higher AUROC values than modified SOFA. A pre-onset-only sensitivity analysis, restricted to patients whose snapshot preceded their first positive sepsis label (achieved median lead time 39 hours), showed all four estimators converging (AUROC 0.579– 0.591) with no significant pairwise differences, indicating the primary-analysis gap over SOFA was partly attributable to postonset records. A stricter analysis limited to a minimum 6-hour lead time reversed the ranking: the SOFA-derived score obtained the highest AUROC (0.597), with the LLMs and NEWS2-derived score converging lower (0.573–0.578), again with no significant pairwise differences. Their alerting behaviour was not interchangeable: GPT-5.5 produced a more balanced sensitivity-specificity profile, while Claude Sonnet 5 identified more positive cases at the cost of additional false alerts. These findings show that zero-shot language models can provide discrimination comparable to the strongest modified clinical score in this dataset. The modest absolute AUROC values, use of modified scores, and retrospective single-dataset design mean that the models should be considered complementary research tools, not stand-alone clinical detectors.

Read PDF

Similar papers

#large language models Open access Sep 2026

Language model-assisted label refinement for accurate sepsis detection from electronic health records

STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality, outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters.

I. K. Kalyvianakis, C. De Amezaga, E. S. Lee et al. · 0 citations
Open access Aug 2026

Early Sepsis Prediction Using Interpretable Models.

Sepsis remains a major cause of morbidity and mortality in Intensive Care Units (ICUs). Timely identification of sepsis can prevent severe complications by enabling early treatment, such as administering antibiotics. Despite advances in diagnostic biomarkers and scoring systems, these approaches often lack the ability...

Charithea Stylianides, Andria Nicolaou, Anna Vavlitou et al. · 0 citations
Review Open access Aug 2026

AI-Driven Predictive Models for Early Detection of Pediatric Sepsis: A Systematic Review and Meta-Analysis

Background Pediatric sepsis continues to pose a major challenge in healthcare, compounded by delayed diagnosis and treatment resulting in poor outcomes. Artificial intelligence (AI) and machine learning (ML) continue to develop predictive models that can support the early identification of pediatric sepsis and assist w...

A. Shibu, J. Aadhira, S. Mitra et al. · 0 citations
Open access Aug 2026

AI-based clinical prediction model for early infectious disease classification in the emergency department

Background Diagnosing infections remains challenging. Clinicians rely on scoring systems and experience, but artificial intelligence (AI) is used to support decision-making by integrating clinical data. However, most AI models focus on predicting adverse outcomes (e.g., ICU admission or sepsis) rather than differentiat...

Sara N. Søgaard, H. Skjøt-Arkil, C. Mogensen et al. · 0 citations
Open access Sep 2026

Evaluating the generalisability of LiSep LSTM for early prediction of septic shock across US and european cohorts

Sepsis remains a major public health concern and is associated with high mortality. Early detection and timely intervention are critical for improving outcomes, yet no standardised approach has been universally adopted. The LiSep LSTM model, a Long Short-Term Memory neural network developed using the MIMIC-III database...

Hong-Yong Tan, L. Froese, D. Wilhelms et al. · 0 citations
Open access Sep 2026

Development of an interpretable machine learning model for predicting in-hospital mortality in ICU patients with sepsis: a retrospective cohort study

An interpretable RF-based machine learning model for predicting in-hospital mortality in patients with sepsis was developed and internally validated and showed acceptable discrimination, reasonable calibration, and potential clinical net benefit in internal validation.

Tian-Yu Zhao, Ke-Xin Wen, Xu-Min Han et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.