Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Dec 5, 2022
Date Accepted: Apr 18, 2023

The final, peer-reviewed published version of this preprint can be found here:

Natural Language Processing for Clinical Laboratory Data Repository Systems: Implementation and Evaluation for Respiratory Viruses

Dolatabadi E, Chen B, Buchan SA, Austin AM, Azimaee M, McGeer A, Mubareka S, Kwong JC

Natural Language Processing for Clinical Laboratory Data Repository Systems: Implementation and Evaluation for Respiratory Viruses

JMIR AI 2023;2:e44835

DOI: 10.2196/44835

PMID: 38875570

PMCID: 11057455

Natural Language Processing for Clinical Laboratory Repository Systems: Implementation and Evaluation for Respiratory Viruses

  • Elham Dolatabadi; 
  • Branson Chen; 
  • Sarah A Buchan; 
  • Alex Marchand Austin; 
  • Mahmoud Azimaee; 
  • Allison McGeer; 
  • Samira Mubareka; 
  • Jeffrey C Kwong

ABSTRACT

Background:

With the growing volume and complexity of laboratory repositories, it has become tedious to parse unstructured data into structured and tabulated formats for secondary uses such as decision support, quality assurance, and outcome analysis. However, advances in Natural Language Processing (NLP) approaches have enabled efficient and automated extraction of clinically meaningful medical concepts from unstructured reports.

Objective:

In this study, we aimed to determine the feasibility of using the NLP model for information extraction as an alternative approach to a time-consuming and operationally resource-intensive handcrafted rule-based tool. Therefore, we sought to develop and evaluate a deep learning-based NLP model to derive knowledge and extract information from text-based laboratory reports sourced from a provincial laboratory repository system.

Methods:

The NLP model, a hierarchical multi-label classifier, was trained on a corpus of laboratory reports covering testing for 14 different respiratory viruses and viral subtypes. The corpus included 85k unique laboratory reports annotated by eight Subject Matter Experts (SME). The model's performance stability and variation were analyzed across fine-grained and coarse-grained classes. Moreover, the model's generalizability was also evaluated internally and externally on various test sets.

Results:

The NLP model was trained several times with random initialization on the development corpus, and the results of the top ten best-performing models are presented in this paper. Overall, the NLP model performed well on internal, out-of-time (pre-COVID-19), and external (different laboratories) test sets with micro-averaged F1 scores >94% across all classes. Higher Precision and Recall scores with less variability were observed for the internal and pre-COVID-19 test sets. As expected, the model's performance varied across categories and virus types due to the imbalanced nature of the corpus and sample sizes per class. There were intrinsically fewer classes of viruses being detected than those tested; therefore, the model's performance (lowest F1-score of 57%) was noticeably lower in the detected cases.

Conclusions:

We demonstrated that deep learning-based NLP models are promising solutions for information extraction from text-based laboratory reports. These approaches enable scalable, timely, and practical access to high-quality and encoded laboratory data if integrated into laboratory information system repositories.


 Citation

Please cite as:

Dolatabadi E, Chen B, Buchan SA, Austin AM, Azimaee M, McGeer A, Mubareka S, Kwong JC

Natural Language Processing for Clinical Laboratory Data Repository Systems: Implementation and Evaluation for Respiratory Viruses

JMIR AI 2023;2:e44835

DOI: 10.2196/44835

PMID: 38875570

PMCID: 11057455

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.