Development and Evaluation of Large Language Model-Assisted Semi-Automated Data Harmonization Pipeline for the Multiple Chronic Disease Disparities Research Consortium: A Proof-of-Concept Study.
Abstract
Background
The increasing availability of machine-readable research data has created a growing need for efficient large-scale data harmonization (DH). Although large language models (LLMs) show promise for reducing the time and labor required for DH, their effective integration into harmonization workflows remains an important methodological challenge.
Objective
To develop and evaluate a semi-automated, human-in-the-loop (HITL) LLM-assisted DH workflow as a proof of concept across eight projects within a research consortium focused on disparities in multiple chronic conditions.
Methods
We developed an LLM-assisted DH workflow that combined preprocessing, semantic mapping, variable response mapping, synthetic data generation, data transformation, and iterative researcher review. Using GPT-4o hosted in a secure academic environment, project-specific survey items were mapped to 114 Consortium Common Data Element (CDE) semantic groups. Semantic mapping accuracy was evaluated based on the Consortium's consensus harmonization results. Complete and incomplete semantic mappings were distinguished according to construct overlap, and incomplete mappings requiring context-dependent human judgement were excluded for mapped item pairs and CDE semantic groups with and without HITL review.
Results
Following preprocessing, 884 survey items were included in the DH workflow, of which 812 were mapped to a mean of 79 CDE semantic groups per project. Mean semantic mapping accuracy for item pairs was 95.9% with HITL review and 86.7% without HITL review. At the CDE semantic-group level, the corresponding accuracies were 98.9% and 94.9%, respectively. Across projects, omission of valid semantic mappings occurred more frequently than incorrect semantic mappings, particularly for heterogeneous response structures and subjective constructs measured using different instruments. Pipeline components-including concept-based preprocessing, iterative prompt refinement, and targeted HITL review-improved mapping completeness and supported variable harmonization across heterogeneous datasets.
Conclusions
This proof-of-concept study demonstrates that the effectiveness of LLM-assisted DH depends more on the design of a structured HITL workflow rather than on the LLM alone. Combining preprocessing, iterative verification, and targeted human oversight improved the efficiency and reliability of DH while preserving researcher judgment for complex mapping decisions. CLINICALTRIAL