Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Oct 13, 2025
Date Accepted: Aug 12, 2026
Symptom Terminology Normalization in Traditional Chinese Medicine: Development and Evaluation of a Two-Stage Deep Learning Framework Based on Fine-Grained Semantic Classification
ABSTRACT
Background:
Due to the heterogeneity of symptom terminology and the lack of industry standards, the same symptom is often described in multiple expressions. Current normalization approaches struggle to comprehensively retrieve standard terms when an input raw term maps to multiple symptoms.
Objective:
To address the lack of industry standards for traditional Chinese medicine (TCM) symptom terminology, this paper proposes the Split-Then-Concatenate Normalization Framework (STC-NF), a novel approach based on fine-grained semantic classification and a two-stage deep learning architecture that utilizes electronic medical records (EMRs) as the data source.
Methods:
This paper proposed a two-stage deep learning framework, "Split-Then-Concatenate". In the Splitting stage, TCM symptom entities were categorized into 12 fine-grained semantic labels, and three Named Entity Recognition (NER) models were trained to extract TCM symptom terminology from EMRs. In the Concatenation stage, standard terms with the same concept as raw terms were identified by a BERT-based Binary Classification model. Then, the standard terms with specific semantic labels were concatenated and reordered according to predefined rules to output structured text, thereby normalizing TCM symptom terminology.
Results:
The accuracy of the proposed model in handling single-implication terms (where one raw term maps to only one standard term) reached 91.48%, an increase of 44.33 percentage points over TF-IDF (47.15%). For multi-implication terms (where one raw term maps to multiple standard terms), the accuracy was 84.20%, outperforming Sequence Generation (51.40%) by 32.80 percentage points. Tested on a mixed dataset of single- and multi-implication terms, the combined accuracy was 88.04%, which was 5.84 percentage points higher than the best-performing baseline model, MTCG (82.20%).
Conclusions:
In this paper, we verified that the fine-grained semantic classification and the two-stage "Split-Then-Concatenate" framework could effectively improve the performance of Named Entity Recognition and Entity Alignment, providing a better normalization approach for TCM symptom terminology.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.