Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Apr 6, 2025
Date Accepted: Apr 24, 2026

The final, peer-reviewed published version of this preprint can be found here:

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

Jang WS, Sultana S, Yao Z, Tran H, Yang Z, Kwon S, Yu H

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

JMIR AI 2026;5:e75561

DOI: 10.2196/75561

PMID: 42467970

Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation: A Comparative Study

  • Won Seok Jang; 
  • Sharmin Sultana; 
  • Zonghai Yao; 
  • Hieu Tran; 
  • Zhichao Yang; 
  • Sunjae Kwon; 
  • Hong Yu

ABSTRACT

Background:

OpenNotes allows patients to access their electronic health record (EHR) notes through online patient portals. However, EHR notes contain abundant medical jargon, which can be difficult for patients to comprehend. One way to improve comprehension is by reducing information overload and helping patients focus on the medical terms that matter most to them.

Objective:

In this study, we evaluated both closed-source and open-source Large Language Models (LLMs) for extracting and prioritizing medical jargon from EHR notes relevant to individual patients, leveraging prompting techniques, fine-tuning and data augmentation.

Methods:

We evaluated the performance of closed-source and open-source LLMs on a dataset of 90 expert-annotated EHR notes. We tested various combinations of settings including: i) general and structured prompts, ii) zero-shot and few-shot prompting, iii) fine-tuning and iv) data augmentation. To enhance the extraction and prioritization capabilities of open-source models in low-resource settings, we applied data augmentation using GPT-4o and integrated a ranking technique to refine the training process. Additionally, to measure the impact of dataset size, we fine-tuned the models by incrementally increasing the size of the augmented dataset from 10 to 9,995 and tested their performance. The effectiveness of the models was assessed using 10-fold cross-validation, providing a comprehensive evaluation across various settings. We report the F1 score and Mean Reciprocal Rank (MRR) for performance evaluation using two different string matching algorithms (relaxed string matching and Jaccard Index). We also conduct error analysis classifying the erroneous outputs from the models.

Results:

Our results show that open-source models achieved the highest performance, particularly when utilizing fine-tuning with gold-standard dataset. Under Jaccard Index based string matching, DeepSeek 8B set the benchmarks with an F1 of 0.431 (SD 0.046); similarly, BioMistral 7B showed an MRR of 0.577 (0.109). However, under relaxed string matching, open-source models were unable to match the performance of closed-source models, even with data augmentation or fine-tuning. We analyzed our experiment from several perspectives. First, few-shot prompting did not show an advantage over zero-shot prompting in vanilla models. Second, when comparing general and structured prompts, we found model performance could deviate largely based on prompting styles. Third, fine-tuning with a small gold truth dataset improved performance. Lastly, data augmentation yielded performance comparable to or even surpassing fine-tuning strategy. However, it also left emphasis on the importance of the quality of the augmented dataset. Conclusion: The evaluation of both closed-source and open-source LLMs highlighted the effectiveness of prompting strategies, fine-tuning, and data augmentation in enhancing model performance in low-resource scenarios.


 Citation

Please cite as:

Jang WS, Sultana S, Yao Z, Tran H, Yang Z, Kwon S, Yu H

Enhancing Large Language Models for Identifying and Prioritizing Important Medical Jargons From Electronic Health Record Notes Using Data Augmentation: Comparative Study

JMIR AI 2026;5:e75561

DOI: 10.2196/75561

PMID: 42467970

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.