Previously submitted to: JMIR AI (no longer under consideration since Jan 09, 2026)
Date Submitted: Jan 13, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge
ABSTRACT
Background:
Large Language Models (LLMs) are increasingly utilized in medical knowledge platforms, yet their capabilities in specialized domains like vaccination remain underexplored. Understanding their performance in these domains is vital for their successful application in educational and public health initiatives.
Objective:
The objective of this study is to evaluate the performance of leading Large Language Models (LLMs) in the specialized domain of vaccination knowledge using a newly developed Vaccination Evaluation Benchmark Dataset. By assessing their accuracy across multiple-choice questions (MCQs) pertaining to 14 common mandatory or recommended vaccines, this study aims to identify the strengths and weaknesses of LLMs in vaccine-related knowledge. Additionally, it seeks to provide insights into the selection and optimization of LLMs for use in public health initiatives and medical education, with a focus on improving their reliability and effectiveness in underrepresented vaccine areas.
Methods:
This study generates the Vaccination Evaluation Benchmark Dataset, comprising 1,341 multiple-choice questions (MCQs) to assess the vaccine knowledge of various leading LLMs, including GPT-4o, Claude3-Opus, and Gemini 1.5 Pro. The dataset draws from global health authoritative sources such as the WHO and CDC for 918 MCQs and peer-reviewed scientific literature from January 2015 to April 2024 for an additional 423 MCQs, covering 14 common mandatory or recommended vaccines such as COVID-19, Dengue, Polio, and RSV.
Results:
The GPT-4o model demonstrated the highest accuracy at 88.6% (95% CI: 87.0-90.3). Other models like Claude 3, GPT-4, and Gemini Pro showed competitive results in the mid-80s, while older models such as GPT-3.5 and Zero-one AI achieved accuracies below 80%. LLMs exhibited robust knowledge in well-publicized vaccine areas like Influenza and COVID-19, yet performed weaker in lesser-known vaccine categories such as RSV and Dengue.
Conclusions:
This study underscores the critical importance of a robust selection process for LLMs in developing digital health tools. Newer models, with more comprehensive training, consistently outperform older ones, particularly in vaccination-related knowledge. To maximize their impact on public health and medical education, selecting the right models and enhancing their training on less-publicized vaccines is essential for creating accurate, reliable, and effective tools in addressing global health challenges.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.