Skip to content
Review Open access

Large Language Model-derived Symptom Clusters and Patient Outcomes in Colorectal Cancer from MIMIC-IV Clinical Notes

Sep 2026 · medRxiv · 0 citations
Medicine

TL;DR

LLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.

Abstract

Background: Prior research on symptom clusters (SCs) in colorectal cancer (CRC) has relied primarily on patient-reported outcome surveys, which capture symptom experience at discrete assessment points rather than the continuous documentation generated during routine care, leaving open whether SCs derived from electronic health record (EHR) text carry the same clinical meaning and predictive value. To construct and validate patient-level symptom co-occurrence networks from large language model (LLM), extracted symptom data in CRC patients, and to test whether resulting SCs predict clinical outcomes. Methods: Using a zero-shot LLM extraction pipeline previously benchmarked against a manually annotated ground truth (Macro F1=0.70 for the best-performing model), we extracted 46 symptoms from 2,728 discharge notes of 1,507 CRC patients in MIMIC-IV. Patient-level symptom co-occurrence networks were constructed independently from Gemini 3.5 Flash and Claude Haiku extractions using phi correlation and Louvain community detection, with sensitivity analyses across correlation thresholds, random seeds, note-aggregation strategy, and bootstrap resampling. Per-cluster symptom burden scores were tested as predictors of in-hospital mortality, 30-day readmission, and 1-year mortality using logistic regression adjusted for age, sex, and (in sensitivity models) metastatic disease. Results: Both LLMs' networks converged on three clinically coherent SCs (Systemic, Gastrointestinal, and CRC Disease-Specific) across sensitivity analyses (cross-model Adjusted Rand Index=0.727; 100-seed Louvain ARI=0.985; bootstrap ARI=0.733). The Systemic Symptom Cluster was the most consistent predictor of in-hospital mortality (OR=1.33) and 1-year mortality (OR=1.41), while the CRC Disease-Specific Cluster specifically and independently predicted 30-day readmission (OR=1.20); both associations were robust to adjustment for metastatic disease. Conclusion: LLM-extracted symptom data recover clinically coherent, reproducible SCs from unstructured discharge notes that carry independent prognostic value for mortality and readmission, supporting the clinical validity of automated, EHR-derived symptom profiling in CRC.

Read PDF

Similar papers

Review Open access Sep 2026

Leveraging Large Language Models for Colorectal Cancer Symptom Extraction from MIMIC-IV Clinical Notes

Background: Much of the symptom burden in colorectal cancer (CRC) patients is documented in unstructured discharge-note narrative, and manual extraction is not scalable. Whether large language models (LLMs) outperform rule-based and named entity recognition (NER) methods has not been rigorously benchmarked. Objective:...

Y. Lee, I. Dinov, X. Hu et al. · 0 citations
#small language model Open access Aug 2026

The Performance of Large Language Models in Extracting Intestinal Symptoms From Electronic Health Records: Retrospective Observational Study

This study provides a systematic comparison of several open-source LLMs on a structured intestinal symptom extraction task and concludes that Qwen3 models offer a favorable balance between accuracy and efficiency, making them suitable for resource-constrained scenarios.

Xin-Yue Zhang, Quan-Yu Wang, Beibei Liu et al. · 0 citations
#machine learning Preprint Sep 2026

Knowledge-Enriched Structured EHR Features for 30-Day Hospital Readmission Prediction on MIMIC-IV

Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods achieve strong performance, they depend on the availability of clinical notes, incur substantial computational costs, and yield representations that lack interpretabilit...

Mohamad Najafi, Hong-Yun Fu, M. Brochhausen et al. · 0 citations
Open access Aug 2026

Evaluation of Diagnostic Accuracy of Open-Source and Proprietary Large Language Models Across Multi-System Clinical Cases

A reproducible estimate of diagnostic retrieval accuracy across four widely used model configurations is provided to establish a baseline for further clinical validation and establish a baseline for further clinical validation.

Lalwani Saurabh, Bodetti Dr.Vishala, Gor Kishan et al. · 0 citations
Review Open access Sep 2026

Practical Guide to Large Language Models for Information Extraction in Behavioral Health Notes: Tutorial

Abstract Background Mental health clinical notes contain decision-critical information often absent from structured electronic health record fields. Large language models (LLMs) can extract clinically relevant signals from narrative text; however, variability in output format, limited reproducibility, and inconsistent...

Diya Saha, J. Edgcomb · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.