Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Formative Research (no longer under consideration since Mar 19, 2026)

Date Submitted: Jan 24, 2026

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Evaluating Large Language Model Adherence to Targeted 5th-Grade Readability Standards in Patient Education on Chronic Conditions: A Cross-Sectional Comparative Study

  • Faheed Shafau; 
  • Chase Wahl; 
  • Garrett Miedema; 
  • Marcus Kado; 
  • Karanfil Ceric; 
  • Najibah Rehman

ABSTRACT

Background:

Most patient education materials in the United States are written at a reading level well above national recommendations, typically averaging between 9th-12th grade, while guidelines recommend a 6th grade level or lower. When patient education materials are written at a reading level that is too high, it can negatively impact patient care, leading to poorer health outcomes.

Objective:

The objective of this study was to evaluate the effect of explicit fifth-grade readability prompting on the linguistic complexity of large language model-generated patient education materials for common chronic conditions.

Methods:

In our study, we collected responses from ChatGPT, Copilot, and Gemini to 12 patient-centered questions regarding common chronic health-conditions, diabetes, cancer, and heart disease (3 chronic conditions x 4 questions each). Additionally, we structured the study into two arms: a prompted arm in which LLMs were instructed to generate responses at a 5th-grade reading level, and an unprompted arm in which no readability instruction was given (3 LLMs x 12 questions, prompted vs. unprompted). Readability scores were calculated using established indices and compared across LLMs and prompt types using linear mixed-effects regression with a random intercept for question.

Results:

Prompting LLMs to generate 5th grade responses produced substantially lower readability grade levels compared with unprompted outputs across all indices. For Flesch-Kincaid, prompted LSMeans ranged from 5.34 to 7.67, whereas unprompted LSMeans ranged from 11.47 to 11.98. Similar patterns were observed for SMOG (prompted 5.37-6.50 vs unprompted 10.60 - 10.88) and Gunning Fog (prompted 6.46 -7.97 vs unprompted 12.18 - 12.84). A significant LLM x Prompt Type interaction was observed for all readability indices (F = 3.34–4.88, p = 0.0417 to 0.0113) indicating that the magnitude of improvement differed across models.

Conclusions:

Prompting LLMs to use a 5th-grade reading level substantially improves the readability of their outputs, though the degree of improvement varies across models. These findings highlight the importance of prompt design and model selection when using LLMs for patient education.


 Citation

Please cite as:

Shafau F, Wahl C, Miedema G, Kado M, Ceric K, Rehman N

Evaluating Large Language Model Adherence to Targeted 5th-Grade Readability Standards in Patient Education on Chronic Conditions: A Cross-Sectional Comparative Study

JMIR Preprints. 24/01/2026:92110

DOI: 10.2196/preprints.92110

URL: https://preprints.jmir.org/preprint/92110

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.