Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Formative Research

Date Submitted: Mar 22, 2026
Date Accepted: Jul 21, 2026

The final, peer-reviewed published version of this preprint can be found here:

Large Language Models for Patient Education in Cardiovascular Imaging: Prospective Observational Comparative Study

Marey A, Pal B, Yaşar AB, Rath S, Francese G, Ghorab H, Niemierko J, Jamill MSW, Umair M

Large Language Models for Patient Education in Cardiovascular Imaging: Prospective Observational Comparative Study

JMIR Form Res 2026;10:e95883

DOI: 10.2196/95883

PMID: 42647860

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Evaluating the Accuracy and Reliability of AI Conversational Agents in Patient Education on Cardiovascular Imaging: An Observational Comparative Study of ChatGPT o1, ChatGPT 4o, and Deepseek

  • Ahmed Marey; 
  • Basudha Pal; 
  • Ayşenur Buz Yaşar; 
  • Shree Rath; 
  • Giulia Francese; 
  • Hossam Ghorab; 
  • Julia Niemierko; 
  • Muhammad Shah Wali Jamill; 
  • Muhammad Umair

ABSTRACT

Background:

Large language models (LLMs) are increasingly used to support digital health communication, yet their reliability in patient-facing cardiovascular imaging education remains uncertain. Cardiovascular imaging involves complex terminology and procedural details that many patients struggle to understand, creating a need for accurate, clear, and reassuring explanations. While prior evaluations of conversational AI have focused primarily on diagnostic reasoning or clinician-oriented tasks, few studies have systematically compared contemporary LLMs in their ability to communicate effectively with patients.

Objective:

To compare the accuracy, clarity, completeness, and patient-centered communication quality of responses generated by three state-of-the-art conversational agents- (i) DeepSeek, (ii) ChatGPT o1, and (iii) ChatGPT 4o when addressing real-world patient questions about cardiovascular imaging.

Methods:

A prospective methodological evaluation was conducted using 84 unique patient-centered questions curated from authoritative cardiovascular information sources and online patient forums. Each question was independently submitted to DeepSeek, ChatGPT o1, and ChatGPT 4o in isolated sessions to avoid contextual contamination. Two cardiovascular radiologists scored each response across four domains: Accuracy, Clarity/Appropriateness, Completeness, and User Engagement/Reassurance, using a standardized 3-point rubric (total score range 4–12). Discrepancies were resolved through predefined adjudication procedures. Descriptive statistics were computed, and group differences were analyzed using one-way ANOVA or Kruskal–Wallis testing as appropriate. Categorical distributions were compared using Pearson’s chi-square test. Statistical significance was defined as α = 0.05 with star-notation thresholds (*p < 0.05, **p < 0.01, ***p < 0.001).

Results:

Across the 84 patient questions, all three models produced largely accurate and complete responses. Mean Accuracy scores were similarly high for DeepSeek (2.82/3), ChatGPT o1 (2.85/3), and ChatGPT 4o (2.82/3). Clarity scores were also comparable for DeepSeek (2.90/3), o1 (2.82/3), and 4o (2.77/3). Completeness showed the same pattern, with scores of 2.79/3, 2.83/3, and 2.83/3, respectively. The only meaningful difference appeared in User Engagement and Reassurance. DeepSeek averaged 2.96/3 and ChatGPT o1 2.99/3, whereas ChatGPT 4o scored markedly lower at 2.54/3. Categorical analysis showed that “good” engagement ratings were assigned to 81/84 DeepSeek responses (96.4%), 83/84 ChatGPT o1 responses (98.8%), but only 45/84 ChatGPT 4o responses (53.6%) with p < 0.001. Total composite scores reflected this pattern: DeepSeek averaged 11.48/12, ChatGPT o1 11.49/12, and ChatGPT 4o 10.96/12. No significant differences were observed for Accuracy (p = 0.325), Clarity (p = 0.119), or Completeness (p = 0.653), and no unsafe statements were identified in any model. Here we define unsafe statements as content that could plausibly lead to patient harm through misinformation, inappropriate reassurance, or deviation from standard cardiovascular imaging practices.

Conclusions:

DeepSeek and ChatGPT o1 consistently delivered accurate, clear, and patient-centered explanations of cardiovascular imaging questions, whereas ChatGPT 4o, despite comparable technical accuracy, provided less engaging and reassuring communication. These findings suggest that affective qualities rather than factual correctness represent the main differentiator among current LLMs in patient-education tasks. As conversational agents become integrated into cardiovascular imaging workflows, attention to communication tone, emotional support, and health-literacy alignment will be essential to ensure safe and effective patient use.


 Citation

Please cite as:

Marey A, Pal B, Yaşar AB, Rath S, Francese G, Ghorab H, Niemierko J, Jamill MSW, Umair M

Large Language Models for Patient Education in Cardiovascular Imaging: Prospective Observational Comparative Study

JMIR Form Res 2026;10:e95883

DOI: 10.2196/95883

PMID: 42647860

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.