Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR AI (no longer under consideration since Jan 09, 2026)

Date Submitted: Jan 13, 2025

VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge

  • Siyuan Chen; 
  • Zhengdong Wu; 
  • Lily Wass; 
  • Lucas Garay; 
  • José Vizoso; 
  • Kathy Leung; 
  • Leesa Lin

Background:

Large language models (LLMs) are increasingly utilized by healthcare professionals and the public seeking health information, yet their accuracy in providing vaccine-related knowledge remains underexplored. Their ability to return correct information comes at a critical time for vaccines, following significant landslides in both vaccine confidence and uptake throughout the COVID-19 pandemic.

Objective:

This study aimed to evaluate the capabilities of advanced LLMs within the specialized domain of vaccination knowledge, leveraging a novel Vaccination Evaluation Benchmark Dataset of validated official or peer-reviewed questions. Through a systematic accuracy assessment of vaccine-related multiple-choice questions (MCQs), this study reveals critical insights into the strengths and limitations of current LLMs’ vaccine knowledge.

Methods:

This study generated a Vaccination Evaluation Benchmark Dataset, covering 14 vaccines in three languages: English, Spanish, and Chinese. The dataset prominently featured widely recognized vaccines such as HPV, COVID-19, and influenza, while also including inquiries on emerging vaccines like RSV and malaria. The dataset was used to evaluate the performance of 13 popular LLMs across three prompting strategies – zero-shot, few-shot, and chain-of-thought (CoT) – with exact-match accuracy of correct answers measured across language and vaccine type. A generalized linear mixed model (GLMM) was constructed to evaluate effects of model release, prompting strategy, question language, topic, and vaccine on correctness of responses.

Results:

The dataset comprised 1,886 MCQs in English (71%), Spanish (13%), and Chinese (16%), focusing primarily on HPV (276), COVID-19 (230), and influenza (192). Most questions addressed dosing and recommendations (41.73%), effectiveness (14.85%), and safety (12.51%), aligning with clinical vaccination priorities. Geographic disparities were evident, with broader vaccine research coverage in North America (7-9 types). In comparison, English-publications from countries in Africa and South America generally focused on only 1-2 different vaccines. Evaluation of 13 LLMs showed accuracies of 86.0%, 83.7%, and 80.0% in English, Spanish, and Chinese, respectively, a pattern confirmed by the GLMM (P < 0.001). Flagship models performed significantly better than older models. In terms of prompting strategy, few-shot prompting yielded the highest accuracy (86%) and had significantly higher odds of generating correct answers compared with zero-shot and CoT strategies. Performance decreased for less widely known vaccines, including Dengue (accuracy 72.7%; Odds Ratio [OR] = 0.29), RSV (67.4%; OR = 0.41), and Shingles (78.7%; OR = 0.42), as well as for some mandatory childhood immunizations, such as Meningococcal (77.3%; OR = 0.42). Questions related to disease awareness and prevention, regulatory monitoring systems, and misconceptions and corrections were answered more accurately than questions in other domains.

Conclusions:

This study emphasizes the necessity of rigorous model selection and targeted refinement in the development and deployment of AI-driven digital health solutions for vaccine confidence. Performance is higher with newer model versions and for questions in English, raising questions about equitable access to high-performing proprietary models as well as potential linguistic and cultural biases in model training.

Clinicaltrial:


 Citation

Please cite as:

Chen S, Wu Z, Wass L, Garay L, Vizoso J, Leung K, Lin L

VaxEval: A Benchmark Dataset for Evaluating Large Language Models on Vaccine Knowledge

DOI: 10.2196/71216

URL: https://preprints.jmir.org/preprint/71216

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.