Skip to content
Conference Open access

Intelligent Narrative Summaries and Risk Scoring of Laboratory Panels with Large Language Models

2026 · EPJ Web of Conferences · 0 citations · 7 references

TL;DR

A prompt-driven pipeline that converts FHIR R4 laboratory panels into structured, paragraph-length clinical narratives paired with a calibrated 0-1 risk score, using GPT-4o-mini as the generation engine is described, demonstrating feasibility for abnormality flagging and narrative generation.

Abstract

Laboratory medicine sits at the intersection of clinical science and data management. A single hospital admission can generate dozens of analyte values, yet most electronic health record (EHR) interfaces present them as rows in a table, leaving interpretation entirely to the clinician. Alert fatigue, driven in part by poorly calibrated notifications remains one of the most documented usability problems in modern EHR design [1]. This paper describes a prompt-driven pipeline that converts FHIR R4 laboratory panels into structured, paragraph-length clinical narratives paired with a calibrated 0-1 risk score, using GPT-4o-mini as the generation engine. The full behavioral specification is encoded in the prompt and output schema. We evaluated the system on 200 laboratory panels, each drawn from a distinct synthetic patient, from Synthea-generated FHIR bundles spanning seven panel categories (metabolic, lipid, blood count, diabetes monitoring, kidney, liver, and urine). We compared three configurations: a rule-only template baseline, the LLM alone (no seed), and the hybrid pipeline in which a deterministic rule-based risk seed is supplied to the LLM. Abnormal-analyte detection was near ceiling and statistically indistinguishable for both LLM configurations (F1 ≈ 0.97), indicating that the model recovers out-of-range analytes directly from the structured table with or without the seed. The seed's measurable contribution is to risk-score calibration: the correlation between the model's 0–1 risk score and the reference rule score rose from r = 0.87 (no seed) to r = 0.96 (with seed). The rule seed thus functions as a calibration mechanism rather than a detection aid. These are proof-of-concept results on synthetic structured data. They demonstrate feasibility for abnormality flagging and narrative generation; they do not constitute a claim of clinical validity, which would require real-world data and clinician review. The paper contributes a reproducible architectural framework, a systematic quantitative benchmark on synthetic panels, and a grounded discussion of the integration challenges and future directions that separate a research prototype from a clinically deployed tool.

Read PDF

Similar papers

Jul 2026

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

CLINLENS is introduced, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms, which exposes a substantial gap between runnable submissions and correct clinical analyses.

Yuan Zhu, Ethan B. Liu, Frank Nie et al. · 0 citations
Review Jul 2026

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis and proposes a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsi...

Jiankang Lu, Panyu Chen, Miriam M. Treggiari et al. · 0 citations
Preprint Aug 2026

MedReaMM: Evaluating Large Multimodal Models on Expert-Level Clinical Diagnostic Synthesis

This work introduces MedReaMM, a benchmark specifically designed to evaluate models'ability to synthesize heterogeneous clinical evidence consisting of detailed patient histories alongside multiple medical images into accurate differential diagnoses under a complete-information paradigm.

Lai Wei, Yu-Chao Chen, Zhenbiao Cao et al. · 0 citations
Open access Sep 2026

Evaluating RAG Configurations for Clinical Information Extraction from EHR Notes: Aged Care Case Study

Clinical information extraction from unstructured electronic health records is important for supporting clinical decision making and healthcare research. However, large language models can struggle to accurately extract domain-specific information without effective adaptation. Retrieval-augmented generation offers a...

Dinithi S. Vithanage, Quang Vinh Duong, Chao Deng et al. · 0 citations
#machine learning Preprint Sep 2026

Knowledge-Enriched Structured EHR Features for 30-Day Hospital Readmission Prediction on MIMIC-IV

Recent approaches to 30-day hospital readmission prediction rely on pre-trained language models applied to discharge summaries. Although these methods achieve strong performance, they depend on the availability of clinical notes, incur substantial computational costs, and yield representations that lack interpretabilit...

Mohamad Najafi, Hong-Yun Fu, M. Brochhausen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.