Skip to content
Review Open access

Towards understanding the disease landscape of clinical trials in Germany: Ontology and embedding-based pipelines versus Large Language Models for ICD-10 Harmonization

Aug 2026 · medRxiv · 0 citations
Medicine

TL;DR

Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses, and indicates that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.

Abstract

Background Clinical trials conducted in Germany are registered across multiple registries, including the German Clinical Trials Register (DRKS), ClinicalTrials.gov, the EU Clinical Trials Register (EUCTR), and, since 2023, the Clinical Trials Information System (CTIS). These registries record health conditions using different classification systems and terminologies, including ICD-10-GM, MeSH, MedDRA, and free text, making cross-registry analyses difficult. We developed and evaluated a pipeline for harmonizing trial condition descriptions to WHO ICD-10 and compared its performance with that of a large language model (LLM) and to health conditions coded by humans. Methods We developed a four-stage, registry-aware mapping pipeline consisting of: (i) condition mention extraction and normalization; (ii) classification of ICD-mappable versus non-mappable mentions; (iii) ontology-based candidate generation using UMLS links between MeSH, MedDRA, ICD-10-GM, and WHO ICD-10; and (iv) SapBERT-based semantic retrieval with hybrid confidence scoring. A second variant additionally applied cross-encoder reranking of the top candidate codes. A stratified sample of 500 condition mentions was manually coded to create an expert reference standard. GPT-4o was evaluated in parallel using the same structured decision framework as the human reviewers. Performance was assessed using accuracy, precision, F1 score, and Cohen's {kappa} at the three-character, block, and chapter levels of ICD-10. Results The pipeline was applied to 23,061 clinical trials and identified 39,512 ICD-mappable condition mentions, of which 72.4% received a high-confidence assignment. Against 390 expert-coded mentions, the baseline pipeline achieved 49.0% accuracy at the three-character ICD-10 level ({kappa} = 0.487), increasing to 58.7% at the chapter level ({kappa} = 0.561). The cross-encoder method produced small but consistent improvements across all evaluation levels. Candidate-recall analysis showed that the correct code was present in the retrieved candidate set in only 73.7% of cases. The LLM substantially outperformed both pipeline variants, achieving 96.7% accuracy and near-perfect agreement with expert coding ({kappa} = 0.966) at the three-character level. The LLM also assigned clinically plausible codes to 82.4% of rejected mentions, 62.8% of Tier-3 exclusions, and 92.3% of review-band mentions. Conclusion Automated harmonization of clinical trial condition data across heterogeneous registries is feasible and supports the use of a common ICD-10 framework for cross-registry analyses. The LLMs achieved high agreement with expert coding, and performed better than the deterministic ontology and embedding pipeline, which achieved moderate agreement. These findings indicate that LLMs can support analyses of the distribution of health conditions investigated in clinical trials in Germany.They are a promising tool for classification of other non-standardised trial characteristics in registries. Keywords: Clinical trial registries; ICD-10; disease harmonization; UMLS; entity linking; SapBERT; large language models; clinical research; natural language processing.

Read PDF

Similar papers

Review Open access Aug 2026

Enhancing healthcare through ontology: a systematic review of challenges and future directions

This study explores the role of ontology in healthcare by surveying numerous research articles to provide a comprehensive overview of its applications, benefits, and challenges. Ontologies, which enable structured representation and integration of complex healthcare knowledge, have been increasingly employed to enhance...

U. Priyadharshini, R. Vijayan · 0 citations
Open access Aug 2026

Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

L. Barrett, N. Joshi, A. S. North et al. · 0 citations
Sep 2026

Abstract B030: From genomic discovery to clinical utility: A framework for evaluating the translational impact of the Gabriella Miller Kids First Program

The National Institutes of Health (NIH) Common Fund’s Gabriella Miller Kids First Pediatric Research Program (Kids First) was established by Congress in 2015 to identify shared genetic pathways between childhood cancer and congenital anomalies. After a decade, NIH leadership sought to understand how Kids First-suppor...

Vanessa I. Barnes · 0 citations
Review Open access Mar 2026

Development and Evaluation of Large Language Model-Assisted Semi-Automated Data Harmonization Pipeline for the Multiple Chronic Disease Disparities Research Consortium: A Proof-of-Concept Study.

BACKGROUND The increasing availability of machine-readable research data has created a growing need for efficient large-scale data harmonization (DH). Although large language models (LLMs) show promise for reducing the time and labor required for DH, their effective integration into harmonization workflows remains an i...

Hyelee Kim, Shuang Liang, K. Lanier et al. · 0 citations
Open access Sep 2026

Real-world drug use in ATC and ICD-10: an expert-curated drug-diagnosis resource based on UK primary care and Danish hospitalization electronic health records

As the use of electronic health records in drug repositioning research increases, so does the need for a well-curated resource describing real-world drug-diagnosis relationships. This need is particularly important in the context of polypharmacy. Although literature-based drug-disease maps exist, they are typically bas...

I. Louloudis, H. Currant, C. Ytsma et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.