Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation
ABSTRACT
Background:
OpenNotes allows patients to access their electronic health record (EHR) notes through online patient portals. However, EHR notes contain abundant medical jargon, which can be difficult for patients to comprehend. One way to improve comprehension is by reducing information overload and helping patients focus on the medical terms that matter most to them.
Objective:
In this study, we evaluated both closed-source and open-source Large Language Models (LLMs) for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning, and data augmentation.
Methods:
We evaluated the performance of closed-source and open-source LLMs on a dataset of 106 expert-annotated EHR notes. We tested various combinations of settings, including: i) general and structured prompts, ii) zero-shot and few-shot prompting, iii) fine-tuning, and iv) data augmentation. To enhance the extraction and prioritization capabilities of open-source models in low-resource settings, we applied data augmentation using ChatGPT and integrated a ranking technique to refine the training process. Additionally, to measure the impact of dataset size, we fine-tuned the models by incrementally increasing the size of the augmented dataset from 10 to 10,000 and tested their performance. The effectiveness of the models was assessed using 5-fold cross-validation, providing a comprehensive evaluation across various settings. We report the F1 score and Mean Reciprocal Rank (MRR) for performance evaluation.
Results:
Among the compared strategies, fine-tuning and data augmentation generally demonstrated higher performance than other approaches. Although the highest F1 score of 0.433 was achieved by GPT-4 Turbo, the highest MRR score of 0.746 was observed with Mistral7B when data augmentation was applied. Notably, by using fine-tuning or data augmentation, open-source models were able to outperform closed-source models. Additionally, achieving the highest F1 score did not always correspond to the highest MRR score. We analyzed our experiment from several perspectives. First, few-shot prompting showed an advantage over zero-shot prompting in vanilla models. Second, when comparing general and structured prompts, each model exhibited different preferences. Third, fine-tuning improved zero-shot performance but sometimes degraded few-shot performance. Lastly, data augmentation yielded performance comparable to or even surpassing that of other strategies.
Conclusions:
The evaluation of both closed-source and open-source LLMs highlighted the effectiveness of prompting strategies, fine-tuning, and data augmentation in enhancing model performance in low-resource scenarios.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.