Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR AI (no longer under consideration since Jan 09, 2026)

Date Submitted: Jan 13, 2025

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge

  • Siyuan Chen; 
  • Zhengdong Wu; 
  • Lily Wass; 
  • Leesa Lin

ABSTRACT

Background:

Large Language Models (LLMs) are increasingly utilized in medical knowledge platforms, yet their capabilities in specialized domains like vaccination remain underexplored. Understanding their performance in these domains is vital for their successful application in educational and public health initiatives.

Objective:

The objective of this study is to evaluate the performance of leading Large Language Models (LLMs) in the specialized domain of vaccination knowledge using a newly developed Vaccination Evaluation Benchmark Dataset. By assessing their accuracy across multiple-choice questions (MCQs) pertaining to 14 common mandatory or recommended vaccines, this study aims to identify the strengths and weaknesses of LLMs in vaccine-related knowledge. Additionally, it seeks to provide insights into the selection and optimization of LLMs for use in public health initiatives and medical education, with a focus on improving their reliability and effectiveness in underrepresented vaccine areas.

Methods:

This study generates the Vaccination Evaluation Benchmark Dataset, comprising 1,341 multiple-choice questions (MCQs) to assess the vaccine knowledge of various leading LLMs, including GPT-4o, Claude3-Opus, and Gemini 1.5 Pro. The dataset draws from global health authoritative sources such as the WHO and CDC for 918 MCQs and peer-reviewed scientific literature from January 2015 to April 2024 for an additional 423 MCQs, covering 14 common mandatory or recommended vaccines such as COVID-19, Dengue, Polio, and RSV.

Results:

The GPT-4o model demonstrated the highest accuracy at 88.6% (95% CI: 87.0-90.3). Other models like Claude 3, GPT-4, and Gemini Pro showed competitive results in the mid-80s, while older models such as GPT-3.5 and Zero-one AI achieved accuracies below 80%. LLMs exhibited robust knowledge in well-publicized vaccine areas like Influenza and COVID-19, yet performed weaker in lesser-known vaccine categories such as RSV and Dengue.

Conclusions:

This study underscores the critical importance of a robust selection process for LLMs in developing digital health tools. Newer models, with more comprehensive training, consistently outperform older ones, particularly in vaccination-related knowledge. To maximize their impact on public health and medical education, selecting the right models and enhancing their training on less-publicized vaccines is essential for creating accurate, reliable, and effective tools in addressing global health challenges.


 Citation

Please cite as:

Chen S, Wu Z, Wass L, Lin L

VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge

JMIR Preprints. 13/01/2025:71216

DOI: 10.2196/preprints.71216

URL: https://preprints.jmir.org/preprint/71216

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.