Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: Dec 18, 2025
Date Accepted: Jul 13, 2026
Date Submitted to PubMed: Jul 15, 2026

The final, peer-reviewed published version of this preprint can be found here:

Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People Living With HIV: Comparative Study

Kong T, Sun L, Luo Y, Xiao X, Li J, Liu J

Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People Living With HIV: Comparative Study

J Med Internet Res 2026;28:e89858

DOI: 10.2196/89858

PMID: 42449481

PMCID: 13456308

Comparative performance of AI models and clinicians in evidence-based cardiovascular disease management for people living with HIV: Comparative Study

  • Tianqi Kong; 
  • Liqin Sun; 
  • Yinsong Luo; 
  • Xi Xiao; 
  • Jin Li; 
  • Jiaye Liu

ABSTRACT

Background:

With increasing life expectancy, cardiovascular disease (CVD) has become a major comorbidity among people living with HIV (PLWH), necessitating multidisciplinary and guideline-driven management. Large language models have shown promise in healthcare, yet their ability to support CVD care in this population remains unclear.

Objective:

This study compared the performance of four mainstream AI models (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH.

Methods:

Twenty-five guideline-based structured questions were developed. AI-generated responses were obtained using standardized prompts, while clinician responses were collected through one-on-one interviews and transcribed verbatim. All responses were independently evaluated by six experts across four dimensions including accuracy, completeness, readability, and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent).

Results:

AI models significantly outperformed clinicians across all dimensions (P < 0.01). AI achieved mean scores of 3.44–3.68 (median 4), reflecting high quality and consistency. By contrast, clinicians had substantially lower mean scores (1.78–2.05; median 2) and much greater variability. Among the AI models, scores differed significantly (F = 24.346, p < 0.001). Post-hoc tests (Tukey-adjusted) showed that DeepSeek-R1 significantly outperformed all others, with score differences of 0.198 (95% CI: 0.134–0.262) against ChatGPT-4o, 0.193 (95% CI: 0.129–0.257) against ChatGPT-o4-mini, and 0.230 (95% CI: 0.166–0.294) against DeepSeek-V3 (all p values < 0.001). Specialty-based comparison showed cardiologists outperformed infectious disease clinicians in CVD risk assessment (2.26 vs. 1.83), whereas infectious disease clinicians performed better in side effects of drugs (2.23 vs. 1.65).

Conclusions:

In this structured Q&A study addressing cardiovascular disease management for people living with HIV, AI models demonstrated superior performance compared to clinicians across all evaluation metrics, particularly the DeepSeek-R1 model which achieved the highest scores. These findings highlight the potential of AI as a decision-support tool to bridge interdisciplinary knowledge gaps and improve the efficiency of clinical communication. To maximize impact, AI should be integrated with multidisciplinary collaboration and targeted clinician training, thereby strengthening the management of complex comorbidities in PLWH.


 Citation

Please cite as:

Kong T, Sun L, Luo Y, Xiao X, Li J, Liu J

Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People Living With HIV: Comparative Study

J Med Internet Res 2026;28:e89858

DOI: 10.2196/89858

PMID: 42449481

PMCID: 13456308

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.