Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Nov 3, 2025
Date Accepted: Jul 8, 2026

The final, peer-reviewed published version of this preprint can be found here:

Measuring Consistency Between Large Language Models’ Responses to Preventive Care Queries and Official US Preventive Services Task Force (USPSTF) Recommendations: Systematic Test Involving All USPSTF Preventive Care Topics via Simulated User Prompts

Johnson T, Gaissmaier W

Measuring Consistency Between Large Language Models’ Responses to Preventive Care Queries and Official US Preventive Services Task Force (USPSTF) Recommendations: Systematic Test Involving All USPSTF Preventive Care Topics via Simulated User Prompts

JMIR AI 2026;5:e87034

DOI: 10.2196/87034

PMID: 42530970

Measuring Consistency between Large Language Models’ Responses to Preventive-Care Queries and Official Recommendations of the U.S. Preventive Services Task Force: A Systematic Test of All USPSTF Preventive-Care Topics via Simulated-User Prompts

  • Tim Johnson; 
  • Wolfgang Gaissmaier

ABSTRACT

Background:

Large language models (LLMs) provide a new source of information about preventive care activities. Existing research has found mixed performance among a small set of LLMs queried about select preventive care activities, thus calling for comprehensive testing of a large set of LLMs on a wider range of preventive care topics.

Objective:

To assess whether a large set of popular LLMs generate output about preventive care consistent with a comprehensive set of recommendations from the U.S. Preventive Services Task Force (USPSTF).

Methods:

We investigated whether 28 popular LLMs produced output consistent with all publicly available USPSTF recommendations and grades (n=142) published as of May 2025. LLMs received queries from simulated users, who asked whether they should participate in particular preventive care activities given their inclusion in a relevant population. LLM raters (κ=0.8893) assessed LLM-USPSTF concordance. Automated methods classified responses to detect sources of LLM-USPSTF discrepancy. The study, then, prompted LLMs to rate preventive care activities for relevant populations using the USPSTF grading scale.

Results:

The LLM with the highest concordance rate generated responses consistent with USPSTF recommendations in 89 of 133 tests (66.92%); the LLM with the lowest rate accorded with USPSTF recommendations in 61 out of 135 tests (45.19%). Automated content analysis indicated that LLM outputs equivocated in 3824 of 3902 tests (98.00%). When prompted to grade preventive care activities for particular populations using the USPSTF grading scale, the highest-performing LLM matched USPSTF grades in 122 out of 142 tests (85.92%); the worst matched in 45 out of 142 tests (31.69%).

Conclusions:

LLM deviation from official guidance indicates the continued importance of directing the public to authoritative sources of preventive care advice. Clinical Trial: N/A


 Citation

Please cite as:

Johnson T, Gaissmaier W

Measuring Consistency Between Large Language Models’ Responses to Preventive Care Queries and Official US Preventive Services Task Force (USPSTF) Recommendations: Systematic Test Involving All USPSTF Preventive Care Topics via Simulated User Prompts

JMIR AI 2026;5:e87034

DOI: 10.2196/87034

PMID: 42530970

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.