Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Jan 30, 2026
Date Accepted: Jul 14, 2026
Large Language Model–Based Agent for Clinical Terminology Standardization Toward Semantic Interoperability: Methodological Study
ABSTRACT
Background:
Semantic interoperability—the ability of disparate health information systems to exchange and consistently interpret clinical data—is a cornerstone of modern digital health, underpinning cross-institutional research, real-world evidence generation, and global health surveillance. Laboratory tests constitute one of the richest clinical data sources, yet multilingual variation and institution-specific naming conventions severely impede their standardized integration across systems.
Objective:
We propose LabBridge, a large language model (LLM) agent designed for multilingual standardization of laboratory tests to Logical Observation Identifiers Names and Codes (LOINC), enabling language-agnostic semantic interoperability without language-specific rules or extensive manual curation.
Methods:
LabBridge integrates linguistic normalization, hybrid retrieval (combining domain-adapted embeddings with LOINC ontology structure), and constrained LLM reasoning within an agentic workflow that enforces terminological consistency and traceability. We evaluated the framework on two large-scale, real-world laboratory datasets—one in Chinese and one in English—representing cross-lingual and cross-institutional heterogeneity. Performance was assessed across five LLMs and compared against a vector-based baseline (BGE-M3) and retrieval-augmented generation (RAG) approaches, using mapping accuracy to a curated reference set of 1,487 clinically relevant LOINC codes as the primary metric.
Results:
At full coverage (Top@100%), LabBridge achieved 83–90% LOINC mapping accuracy across all tested LLMs on both languages, substantially outperforming baseline methods. On the Chinese dataset, it improved over BGE-M3 by +66 percentage points (90% vs. 24%) and over the best RAG method by +41 points (90% vs. 49%). On the English dataset, gains ranged from +4 to +27 points over RAG baselines. The framework maintained robust performance across frequency strata, including low-occurrence tests (Top@80–100%), demonstrating resilience to long-tailed term distributions. The highest accuracy—90% on both languages—was achieved using DeepSeek-V3, with GPT-4o (88–89%) and GPT-4o-mini (87–90%) showing comparable results.
Conclusions:
LabBridge demonstrates that embedding LLMs within an ontology-aware, agent-coordinated architecture enables high-fidelity multilingual standardization of laboratory data without language-specific rules or manual curation. By unifying semantic retrieval, linguistic normalization, and constrained reasoning, the framework provides a scalable and reproducible solution for achieving semantic interoperability in diverse, multilingual healthcare environments.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.