Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Jan 30, 2026
Date Accepted: Jul 14, 2026

The final, peer-reviewed published version of this preprint can be found here:

Clinical Laboratory Terminology Standardization for Semantic Interoperability Using a Large Language Model–Based Agent: Methodological Study

Wu L, Huang J, Wang H, Mai L, Zhan X, He X, Zhang X, Liang H, Li X, Bellou A

Clinical Laboratory Terminology Standardization for Semantic Interoperability Using a Large Language Model–Based Agent: Methodological Study

J Med Internet Res 2026;28:e92499

DOI: 10.2196/92499

PMID: 42623484

Large Language Model–Based Agent for Clinical Terminology Standardization Toward Semantic Interoperability: Methodological Study

  • Lijuan Wu; 
  • Jinxin Huang; 
  • Hongnian Wang; 
  • Liyi Mai; 
  • Xueyun Zhan; 
  • Xinrong He; 
  • Xiaotang Zhang; 
  • Huiying Liang; 
  • Xin Li; 
  • Abdelouahab Bellou

ABSTRACT

Background:

Semantic interoperability—the ability of disparate health information systems to exchange and consistently interpret clinical data—is a cornerstone of modern digital health, underpinning cross-institutional research, real-world evidence generation, and global health surveillance. Laboratory tests constitute one of the richest clinical data sources, yet multilingual variation and institution-specific naming conventions severely impede their standardized integration across systems.

Objective:

We propose LabBridge, a large language model (LLM) agent designed for multilingual standardization of laboratory tests to Logical Observation Identifiers Names and Codes (LOINC), enabling language-agnostic semantic interoperability without language-specific rules or extensive manual curation.

Methods:

LabBridge integrates linguistic normalization, hybrid retrieval (combining domain-adapted embeddings with LOINC ontology structure), and constrained LLM reasoning within an agentic workflow that enforces terminological consistency and traceability. We evaluated the framework on two large-scale, real-world laboratory datasets—one in Chinese and one in English—representing cross-lingual and cross-institutional heterogeneity. Performance was assessed across five LLMs and compared against a vector-based baseline (BGE-M3) and retrieval-augmented generation (RAG) approaches, using mapping accuracy to a curated reference set of 1,487 clinically relevant LOINC codes as the primary metric.

Results:

At full coverage (Top@100%), LabBridge achieved 83–90% LOINC mapping accuracy across all tested LLMs on both languages, substantially outperforming baseline methods. On the Chinese dataset, it improved over BGE-M3 by +66 percentage points (90% vs. 24%) and over the best RAG method by +41 points (90% vs. 49%). On the English dataset, gains ranged from +4 to +27 points over RAG baselines. The framework maintained robust performance across frequency strata, including low-occurrence tests (Top@80–100%), demonstrating resilience to long-tailed term distributions. The highest accuracy—90% on both languages—was achieved using DeepSeek-V3, with GPT-4o (88–89%) and GPT-4o-mini (87–90%) showing comparable results.

Conclusions:

LabBridge demonstrates that embedding LLMs within an ontology-aware, agent-coordinated architecture enables high-fidelity multilingual standardization of laboratory data without language-specific rules or manual curation. By unifying semantic retrieval, linguistic normalization, and constrained reasoning, the framework provides a scalable and reproducible solution for achieving semantic interoperability in diverse, multilingual healthcare environments.


 Citation

Please cite as:

Wu L, Huang J, Wang H, Mai L, Zhan X, He X, Zhang X, Liang H, Li X, Bellou A

Clinical Laboratory Terminology Standardization for Semantic Interoperability Using a Large Language Model–Based Agent: Methodological Study

J Med Internet Res 2026;28:e92499

DOI: 10.2196/92499

PMID: 42623484

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.