Previously submitted to: JMIR AI (no longer under consideration since Jan 09, 2026)
Date Submitted: Jan 13, 2025
VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge
Background:
Large language models (LLMs) are increasingly utilized by healthcare professionals and the public seeking health information, yet their accuracy in providing vaccine-related knowledge remains underexplored. Their ability to return correct information comes at a critical time for vaccines, following significant landslides in both vaccine confidence and uptake throughout the COVID-19 pandemic.
Objective:
This study aimed to evaluate the capabilities of advanced LLMs within the specialized domain of vaccination knowledge, leveraging a novel Vaccination Evaluation Benchmark Dataset of validated official or peer-reviewed questions. Through a systematic accuracy assessment of vaccine-related multiple-choice questions (MCQs), this study reveals critical insights into the strengths and limitations of current LLMs’ vaccine knowledge.
Methods:
This study generated a Vaccination Evaluation Benchmark Dataset, covering 14 vaccines in three languages: English, Spanish, and Chinese. The dataset prominently featured widely recognized vaccines such as HPV, COVID-19, and influenza, while also including inquiries on emerging vaccines like RSV and malaria. The dataset was used to evaluate the performance of 13 popular LLMs across three prompting strategies – zero-shot, few-shot, and chain-of-thought (CoT) – with exact-match accuracy of correct answers measured across language and vaccine type. A generalized linear mixed model (GLMM) was constructed to evaluate effects of model release, prompting strategy, question language, topic, and vaccine on correctness of responses.
Results:
The dataset comprised 1,886 MCQs in English (71%), Spanish (13%), and Chinese (16%), focusing primarily on HPV (276), COVID-19 (230), and influenza (192). Most questions addressed dosing and recommendations (41.73%), effectiveness (14.85%), and safety (12.51%), aligning with clinical vaccination priorities. Geographic disparities were evident, with broader vaccine research coverage in North America (7-9 types). In comparison, English-publications from countries in Africa and South America generally focused on only 1-2 different vaccines. Evaluation of 13 LLMs showed accuracies of 86.0%, 83.7%, and 80.0% in English, Spanish, and Chinese, respectively, a pattern confirmed by the GLMM (P < 0.001). Flagship models performed significantly better than older models. In terms of prompting strategy, few-shot prompting yielded the highest accuracy (86%) and had significantly higher odds of generating correct answers compared with zero-shot and CoT strategies. Performance decreased for less widely known vaccines, including Dengue (accuracy 72.7%; Odds Ratio [OR] = 0.29), RSV (67.4%; OR = 0.41), and Shingles (78.7%; OR = 0.42), as well as for some mandatory childhood immunizations, such as Meningococcal (77.3%; OR = 0.42). Questions related to disease awareness and prevention, regulatory monitoring systems, and misconceptions and corrections were answered more accurately than questions in other domains.
Conclusions:
This study emphasizes the necessity of rigorous model selection and targeted refinement in the development and deployment of AI-driven digital health solutions for vaccine confidence. Performance is higher with newer model versions and for questions in English, raising questions about equitable access to high-performing proprietary models as well as potential linguistic and cultural biases in model training.
Clinicaltrial:
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.