Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Dec 18, 2025
Date Accepted: Jul 13, 2026
Date Submitted to PubMed: Jul 15, 2026
Comparative performance of AI models and clinicians in evidence-based cardiovascular disease management for people living with HIV: Comparative Study
ABSTRACT
Background:
With increasing life expectancy, cardiovascular disease (CVD) has become a major comorbidity among people living with HIV (PLWH), necessitating multidisciplinary and guideline-driven management. Large language models have shown promise in healthcare, yet their ability to support CVD care in this population remains unclear.
Objective:
This study compared the performance of four mainstream AI models (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for PLWH.
Methods:
Twenty-five guideline-based structured questions were developed. AI-generated responses were obtained using standardized prompts, while clinician responses were collected through one-on-one interviews and transcribed verbatim. All responses were independently evaluated by six experts across four dimensions including accuracy, completeness, readability, and reliability, using a 4-point ordinal scale ranging from 1 (poor) to 4 (excellent).
Results:
AI models significantly outperformed clinicians across all dimensions (P < 0.01). AI achieved mean scores of 3.44–3.68 (median 4), reflecting high quality and consistency. By contrast, clinicians had substantially lower mean scores (1.78–2.05; median 2) and much greater variability. Among the AI models, scores differed significantly (F = 24.346, p < 0.001). Post-hoc tests (Tukey-adjusted) showed that DeepSeek-R1 significantly outperformed all others, with score differences of 0.198 (95% CI: 0.134–0.262) against ChatGPT-4o, 0.193 (95% CI: 0.129–0.257) against ChatGPT-o4-mini, and 0.230 (95% CI: 0.166–0.294) against DeepSeek-V3 (all p values < 0.001). Specialty-based comparison showed cardiologists outperformed infectious disease clinicians in CVD risk assessment (2.26 vs. 1.83), whereas infectious disease clinicians performed better in side effects of drugs (2.23 vs. 1.65).
Conclusions:
In this structured Q&A study addressing cardiovascular disease management for people living with HIV, AI models demonstrated superior performance compared to clinicians across all evaluation metrics, particularly the DeepSeek-R1 model which achieved the highest scores. These findings highlight the potential of AI as a decision-support tool to bridge interdisciplinary knowledge gaps and improve the efficiency of clinical communication. To maximize impact, AI should be integrated with multidisciplinary collaboration and targeted clinician training, thereby strengthening the management of complex comorbidities in PLWH.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.