Skip to content
Review Open access

Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

LLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.

Abstract

Background: Prior research on symptom clusters (SCs) in colorectal cancer (CRC) has relied primarily on patient-reported outcome surveys, which capture symptom experience at discrete assessment points rather than the continuous documentation generated during routine care, leaving open whether SCs derived from electronic health record (EHR) text carry the same clinical meaning and predictive value. To construct and validate patient-level symptom co-occurrence networks from large language model (LLM), extracted symptom data in CRC patients, and to test whether resulting SCs predict clinical outcomes. Methods: Using a zero-shot LLM extraction pipeline previously benchmarked against a manually annotated ground truth (Macro F1=0.70 for the best-performing model), we extracted 46 symptoms from 2,728 discharge notes of 1,507 CRC patients in MIMIC-IV. Patient-level symptom co-occurrence networks were constructed independently from Gemini 3.5 Flash and Claude Haiku extractions using phi correlation and Louvain community detection, with sensitivity analyses across correlation thresholds, random seeds, note-aggregation strategy, and bootstrap resampling. Per-cluster symptom burden scores were tested as predictors of in-hospital mortality, 30-day readmission, and 1-year mortality using logistic regression adjusted for age, sex, and (in sensitivity models) metastatic disease. Results: Both LLMs' networks converged on three clinically coherent SCs (Systemic, Gastrointestinal, and CRC Disease-Specific) across sensitivity analyses (cross-model Adjusted Rand Index=0.727; 100-seed Louvain ARI=0.985; bootstrap ARI=0.733). The Systemic Symptom Cluster was the most consistent predictor of in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41), while the CRC Disease-Specific Cluster specifically and independently predicted 30-day readmission (OR=1.20); both associations were robust to adjustment for metastatic disease. Conclusion: LLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.

Read PDF

Similar papers

#small language model Open access Aug 2026

The Performance of Large Language Models in Extracting Intestinal Symptoms From Electronic Health Records: Retrospective Observational Study

This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.

Xin-Yue Zhang, Quan-Yu Wang, Beibei Liu et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.