Skip to content

Language model-assisted label refinement for accurate sepsis detection from electronic health records

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality, outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters.

Abstract

Sepsis is a leading cause of hospital mortality, yet timely recognition is hampered by nonspecific presentations and label noise in code-based case definitions. We developed STRIDE, a machine-learning framework for sepsis detection across seven hospitals with a scalable approach to label quality. We refined a pragmatic operational definition using a large language model applied to discharge summaries, with an independent physician-adjudicated cohort as the gold standard. We compared 8-, 24-, and 48-hour observation windows and benchmarked against SOFA, SIRS, and Epic, assessing calibration and discrimination. Among 356,610 encounters, the 8-hour model achieved an AUC of 0.960 in derivation and 0.878 in physician-adjudicated validation, matching or outperforming longer-window models. STRIDE outperformed SOFA and Epic on discrimination and showed favorable calibration by Brier score, while retaining strong discrimination among SIRS-positive non-septic encounters and reaching 78.2% specificity at 80% sensitivity in validation. These findings support accurate sepsis surveillance while limiting unnecessary alerts and requiring minimal prior history.

Read PDF

Similar papers

Review Open access Sep 2026

Beyond ICD Codes: Fine-Tuning LLMs for In-Hospital Cardiac Arrest Identification from EHR Notes

Background: Identification of in-hospital cardiac arrest (IHCA) through manual chart abstraction is the gold standard, but its time-consuming and resource-intensive, which naturally limits its practicality. Diagnostic codes provide an accessible automated alternative, but prior work has shown this method of extraction...

D. Weissenbacher, J. Vo, A. Uy-Evanado et al. · 0 citations
Open access Aug 2026

Early Sepsis Prediction Using Interpretable Models.

Sepsis remains a major cause of morbidity and mortality in Intensive Care Units (ICUs). Timely identification of sepsis can prevent severe complications by enabling early treatment, such as administering antibiotics. Despite advances in diagnostic biomarkers and scoring systems, these approaches often lack the ability...

Charithea Stylianides, Andria Nicolaou, Anna Vavlitou et al. · 0 citations
Open access Sep 2026

Large language models for pneumonia detection in radiology reports via text analysis

Pneumonia is a common infection in critically ill patients with poor outcomes, and its identification relies heavily on radiological evidence. Most MIMIC-based pneumonia studies rely on structured codes rather than radiological text, which may limit the precision of pneumonia-related case identification. This study eva...

Wei-Sheng Chen, Xin-Ya Li, Zhong-Kai Qu et al. · 0 citations
Open access Sep 2026

HeartVar: An LLM-Assisted Tool for Clinical Classification of Variants in Cardiovascular Disease Cohorts

Manual clinical DNA variant classification is the bottleneck of every clinical and research rare disease workflow. The process typically requires a curator to assemble evidence from numerous databases, weigh 28 criteria, reconcile competing evidence, and produce a defensible case for the final classification. Additiona...

Jamie-Lee M. Thompson, Debjani Das, Sally L. Dunwoodie et al. · 0 citations

Related blog posts

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.