Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models
Abstract
Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German emergency-department reports were manually annotated for nausea, vomiting, diarrhea, and dysuria. After reserving 100 reports for prompt formulation and temperature testing, nine LLMs were screened on symptom-specific stratified development sets (N = 250). Selected models were compared with a negation-aware rule-based baseline in independent validation sets (N = 750) using F1-score and patient-level bootstrap confidence intervals. Temperature 0.0 provided the greatest overall stability. Validation F1-scores were 0.985 for vomiting, 0.979 for nausea, 0.824 for dysuria, and 0.814 for diarrhea. Corresponding baseline F1-scores were 0.913, 0.724, 0.705, and 0.853, respectively. Paired comparisons favored LLMs for nausea and vomiting; confidence intervals included zero for diarrhea and dysuria. Median inference times ranged from 0.298 to 1.653 s per report. Discrepancies reflected operational criteria, temporal variation, inconsistent documentation, missed mentions, and five reference errors. Pragmatic model screening can identify suitable local LLMs for clinical free-text annotation. The pipeline is reproducible and adaptable but requires context-specific configuration and validation.