Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Sep 15, 2025
Date Accepted: Jul 8, 2026

The final, peer-reviewed published version of this preprint can be found here:

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

Ma J, Wu X, Feng Z, Kuang Y, Ding Z, Li M, Yang G

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

J Med Internet Res 2026;28:e84086

DOI: 10.2196/84086

PMID: 42574561

Intelligent Framework for ADE Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

  • Junlong Ma; 
  • Xuehong Wu; 
  • Zeying Feng; 
  • Yun Kuang; 
  • Zhendong Ding; 
  • Min Li; 
  • Guoping Yang

ABSTRACT

Background:

Extracting adverse drug event (ADE) information from unstructured clinical notes is challenging due to the semantic complexity. While large language models (LLMs) show promise, they are prone to “hallucinations”.

Objective:

This study aims to evaluate the effectiveness of retrieval-augmented generation (RAG) in improving the identification of ADEs using LLMs from Chinese clinical narratives and to establish a paradigm for this task.

Methods:

We collected and preprocessed 19,983 Chinese clinical notes, from which we created a gold-standard reference set (n=2,510) and an ADE knowledge base (n=5,144). Three LLMs (DeepSeek-V3, Ernie 3.5-8K, and GPT-4o) were evaluated using three distinct prompt strategies: none-augmented generation (NAG), static-augmented generation (SAG), and RAG. Performance was rigorously assessed using precision, recall, and F1 score via 10-fold cross-validation.

Results:

We successfully constructed and publicly released the first Chinese ADE corpus from clinical notes. The RAG framework significantly improved the performance across all tested LLMs compared to NAG and SAG. The optimal configuration, DeepSeek-V3 with RAG, achieved an impressive overall F1 score of 0.9614. Notably, RAG dramatically increased the recall rate for models such as GPT-4o (from 64.76% to 92.39%) and effectively resolved common identification errors that were intractable for non-augmented models.

Conclusions:

Integrating a curated knowledge base with LLMs via a RAG framework is a highly effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes. This approach provides a robust technical foundation for pharmacovigilance and holds significant potential to enhance drug safety research and clinical practice.


 Citation

Please cite as:

Ma J, Wu X, Feng Z, Kuang Y, Ding Z, Li M, Yang G

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study

J Med Internet Res 2026;28:e84086

DOI: 10.2196/84086

PMID: 42574561

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.