Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Oct 29, 2025
Date Accepted: Aug 18, 2026
Comparison of two AI chatbots for diagnosis and providing treatment suggestions in retinopathy of prematurity: A retrospective study
ABSTRACT
Background:
Background:
Retinopathy of prematurity (ROP) is a leading cause of preventable childhood blindness, yet a global shortage of experienced pediatric ophthalmologists impedes timely diagnosis and treatment. While emerging artificial intelligence (AI) chatbots are promising clinical decision-support tools in some ophthalmic diseases, their capabilities in ROP diagnosis and providing treatment suggestions remains uncertain.
Objective:
Purpose: To compare the capabilities of Google’s Gemini 2.5 Pro and OpenAI’s ChatGPT o4-mini in the diagnosis and providing treatment suggestions for retinopathy of prematurity.
Methods:
Methods:
A retrospective analysis was conducted on 70 infants with treatment-requiring ROP, with each infant including structured clinical text data and wide-field fundus images. We adopted a two-stage prompting strategy for AI chatbots, instructing them first to generate ROP diagnoses (including zone, stage, and presence of plus disease) and subsequently to provide treatment suggestions. After collecting the generated responses, we assessed their capabilities by comparing the consistency of their diagnosis and treatment suggestions with the consensus of gold standard. Furthermore, outputs from Gemini 2.5 pro and ChatGPT o4-mini were quantitatively assessed using the ROP-specific Global Quality Score (GQS), graded on a five-point scale from 1 (poor) to 5 (excellent). Statistical significance was determined at the P < 0.05, with all statistical analyses were performed using R software (version 4.4.1; R Foundation for Statistical Computing).
Results:
Results:
For the tasks of ROP zoning, staging, and treatment requirement, there was no statistically significant difference between Gemini 2.5 Pro and ChatGPT o4-mini. However, Gemini 2.5 Pro exhibited statistically superior performance in identifying plus disease (P < .001). In contrast, ChatGPT o4-mini outperformed Gemini 2.5 Pro in providing treatment suggestions, with significantly higher GQS scores in diagnosis (P < .001) and treatment suggestions (P = 0.03).
Conclusions:
Conclusion: ChatGPT o4-mini is a more promising candidate for a general-purpose clinical decision support tool, while Gemini 2.5 Pro’s visual acuity highlights its potential in targeted diagnostic screening.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.