Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: JMIR AI

Date Submitted: Mar 9, 2026
Date Accepted: Aug 31, 2026
Date Submitted to PubMed: Sep 1, 2026

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/93996

The final accepted version (not copyedited yet) is in this tab.

An "ahead-of-print" version has been submitted to Pubmed, see PMID: 42680206

Development and Evaluation of Large Language Model-Assisted Semi-Automated Data Harmonization Pipeline for the Multiple Chronic Disease Disparities Research Consortium: A Proof-of-Concept Study

  • Hyelee Kim; 
  • Shuang Liang; 
  • Kathy Lanier; 
  • Sarit Helman; 
  • Nancy Fan Cheng; 
  • Niloufar Ameli; 
  • Jaysón Davidson; 
  • Shivani Mehta; 
  • Yulin Hswen; 
  • Catherine E. Oldenburg; 
  • Kim Rhoads; 
  • Edwin D. Charlebois; 
  • Stuart A. Gansky; 
  • William Brown III

ABSTRACT

Background:

The increasing availability of machine-readable research data has created a growing need for efficient large-scale data harmonization (DH). Although large language models (LLMs) show promise for reducing the time and labor required for DH, effective integration into harmonization workflows remains an important methodological challenge.

Objective:

To develop and evaluate a semi-automated, human-in-the-loop (HITL) LLM-assisted DH workflow as a proof of concept across eight projects within a research consortium focused on disparities in multiple chronic conditions.

Methods:

We developed an LLM-assisted DH workflow that combined preprocessing, semantic mapping, variable response mapping, synthetic data generation, data transformation, and iterative researcher review. Using GPT-4o hosted in a secure academic environment, project-specific survey items were mapped to 114 Consortium Common Data Element (CDE) semantic groups. Semantic mapping accuracy was evaluated based on the Consortium’s consensus harmonization results. Complete and incomplete semantic mappings were distinguished according to construct overlap, and incomplete mappings requiring context-dependent human judgement were excluded for mapped item pairs and CDE semantic groups with and without HITL review.

Results:

Following preprocessing, 884 survey items were included in the DH workflow, of which 812 were mapped to a mean of 79 CDE semantic groups per project. Mean semantic mapping accuracy for item pairs was 95.9% with HITL review and 86.7% without HITL review. At the CDE semantic-group level, the corresponding accuracies were 98.9% and 94.9%, respectively. Across projects, omission of valid semantic mappings occurred more frequently than incorrect semantic mappings, particularly for heterogeneous response structures and subjective constructs measured using different instruments. Pipeline components—including concept-based preprocessing, iterative prompt refinement, and targeted HITL review—improved mapping completeness and supported variable harmonization across heterogeneous datasets.

Conclusions:

This proof-of-concept study demonstrates an effective LLM-assisted DH depends on the design of a structured HITL workflow rather than on the LLM alone. Combining preprocessing, iterative verification, and targeted human oversight improved the efficiency and reliability of DH while preserving researcher judgment for complex mapping decisions.


 Citation

Please cite as:

Kim H, Liang S, Lanier K, Helman S, Cheng NF, Ameli N, Davidson J, Mehta S, Hswen Y, Oldenburg CE, Rhoads K, Charlebois ED, Gansky SA, Brown W III

Development and Evaluation of Large Language Model-Assisted Semi-Automated Data Harmonization Pipeline for the Multiple Chronic Disease Disparities Research Consortium: A Proof-of-Concept Study

JMIR AI. 31/08/2026:93996 (forthcoming/in press)

DOI: 10.2196/93996

URL: https://preprints.jmir.org/preprint/93996

PMID: 42680206

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.