Accepted for/Published in: JMIR Medical Informatics
Date Submitted: Jan 19, 2026
Date Accepted: Aug 10, 2026
Detecting Misspelled Drug Names Using Transformer-based Language Models
ABSTRACT
Background:
Misspellings in medication names can compromise patient safety, reduce data utility, and impede large-scale data initiatives that integrate medication information from electronic health records (EHRs). Existing methods for detecting misspelled medical terms are mostly dictionary-based and can lead to high false-positive rates when correctly spelled but previously unseen (out-of-vocabulary) terms are encountered.
Objective:
We aimed to develop and validate domain-specific, transformer-based language models for detecting misspelled drug names, with an emphasis on performance for unseen terms.
Methods:
Using RxNorm as a standardized drug vocabulary, we created an RxNorm-augmented training corpus and developed two Bidirectional Encoder Representations from Transformers (BERT)-based modelsāBERTDrug and CharBERTDrugāfor misspelling detection. Specifically, we randomly split 69,824 RxNorm drug names into training, development, and test sets (3:1:1) and generated k misspellings per name using text-perturbation techniques (k optimized for training; fixed at 1 for development and test sets). The models were fine-tuned on the training and development sets and evaluated using the RxNorm test set and 3,586 drug names from the Long-Term Care Data Cooperative (LTCDC) database (external validation). The RxNorm test set and out-of-vocabulary LTCDC dataset (1,922 terms), neither overlapping with the RxNorm training data, were used to evaluate performance on unseen terms. SpellChecker served as a dictionary-based baseline, while fastTextMLand BioWordVecML, which used different subword embeddings as inputs for machine learning, served as additional baselines. Additionally, we compared model performance with GPT-4o, a generative large language model (LLM), using 2,200 randomly sampled test terms.
Results:
BERTDrug and CharBERTDrug outperformed the baseline models on the RxNorm test set on most performance metrics, with BERTDrug performed best (F1=0.852, 95% CI: [0.848, 0.856]; ROC-AUC=0.941, 95% CI: [0.938, 0.943]). In addition, both models outperformed the baseline models on the out-of-vocabulary LTCDC dataset on most metrics, with CharBERTDrug performed best (F1=0.696, 95% CI: [0.676, 0.718]; ROC-AUC=0.788, 95% CI: [0.767, 0.808]). Both models also exceeded GPT-4o on most metrics (except Recall) for RxNorm terms. BERTDrug performed best (0.951 ROC-AUC; 0.855 F1), followed by CharBERTDrug (0.911 ROC-AUC; 0.831 F1) and GPT-4o (0.856 ROC-AUC; 0.721 F1). In contrast, for LTCDC terms, GPT-4o achieved better performance on most metrics, except precision and specificity.
Conclusions:
Domain-specific language models improved detection of misspellings in out-of-vocabulary drug names and outperformed baseline models in both internal and external evaluations. The comparison with a generative LLM suggests that domain shift may substantially reduce the advantages conferred by domain-specific training. With further fine-tuning on diverse data that capture the terminology, formatting conventions, and spelling patterns encountered across real-world clinical settings, these models could be adapted for use in other clinical databases and EHR systems to improve medication data quality for research and to support future safety-focused applications.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.