Skip to content
Open access

Clinical Code Mapping with LLM Tool Use: A Pilot for Automated Data Extraction of Medication and Diagnosis Information from Unstructured Clinical Notes.

Sep 2026 · Studies in Health Technology and Informatics · Vol 340, pp. 172-178 · 0 citations
Medicine

TL;DR

LLMs are suitable for information extraction of medications from clinical notes for use in research databases, however, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.

Abstract

INTRODUCTION Accurate clinical coding is fundamental to large-scale epidemiological studies, hospital billing, and the development of robust clinical decision support systems. Conventional methods for structured data extraction often rely on manual curation, which is prohibitively labor-intensive. Goal of this project is to determine whether current state-of-the-art open-weight LLM models are suitable for extraction of structured data from non-English (German) clinical notes.

Methods

We anonymized 35 German doctor's notes of five patients from our hospital and developed one pipeline to extract and map medications and two for diagnoses. The latter compares a RAG based approach with an agentic AI. We ran these using three open-weight LLMs on a local GPU-PC.

Results

The F1 scores for diagnoses do not exceed 0.12. If we instead consider mapping to the broad category, then the F1 score increases to 0.18. For medications, the F1 score is as high as 0.78 and even 0.95 if we consider trivial name extraction only.

Discussion

For trivial name extraction of medications, every encountered mistake is explainable. Due to limitations in the nature of the task, it is infeasible to expect a perfect score of 1 in any of the coding scenarios. Further problems in LLM output and parsing are addressed.

Conclusion

LLMs excel at extraction. They are suitable for information extraction of medications from clinical notes for use in research databases. However, for a clinical setting where the treatment of patients would be dependent on LLM performance, the current state-of-the-art open weight models are not accurate enough.

Read PDF

Similar papers

Review Open access Sep 2026

Augmenting structured diagnoses through effective use of pre-trained large language models on clinical notes

Abstract Objective Clinical narrative provides a unique window into provider reasoning and attribution for automated diagnosis assignment, but large language models (LLMs) have traditionally not performed well at medical coding. We evaluate a reproducible method for automated diagnosis assignment using LLMs in clinical...

H. Razzaghi, Nhat Nguyen, M. Pargi et al. · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Open access Sep 2026

Development and validation of a pragmatic pipeline for clinical free-text annotation using locally deployed open-weight large language models

Clinical information required for surgical data science (SDS) is frequently embedded in unstructured text. We developed and evaluated a reproducible pipeline for selecting locally deployed open-weight large language models (LLMs) for binary symptom annotation. In this retrospective single-center study, 1,100 German eme...

Jonas Henn, Alisa Stoll, P. Feodorovici et al. · 0 citations
#large language models Review Open access Sep 2026

Large Language Models for Clinical Note Simplification: A Systematic Review and Experimental Evaluation of Medical Text Readability.

Since the introduction of the Patient Rights Act, patients in Germany have gained legal access to their medical records, including clinical notes. However, these documents are typically written for healthcare professionals and are often difficult for patients to understand due to specialized terminology, abbreviations,...

M. Teichmann, Pelin Özkara Menekseoglu, Julian Schwarz et al. · 0 citations
Open access Sep 2026

HeartVar: An LLM-Assisted Tool for Clinical Classification of Variants in Cardiovascular Disease Cohorts

Manual clinical DNA variant classification is the bottleneck of every clinical and research rare disease workflow. The process typically requires a curator to assemble evidence from numerous databases, weigh 28 criteria, reconcile competing evidence, and produce a defensible case for the final classification. Additiona...

Jamie-Lee M. Thompson, Debjani Das, Sally L. Dunwoodie et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.