Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Assessment of ChatGPT and Google in delivering gynecologic cancer education content: A comparative evaluation of accuracy, completeness, and reference quality
ABSTRACT
Background:
Patients with newly diagnosed gynecologic (GYN) cancers often seek information online, but the quality of available resources may be inconsistent. While ChatGPT may offer an alternative to traditional internet search engines, its performance remains largely understudied in GYN oncology.
Objective:
To compare the completeness, accuracy, and reference quality of patient education responses generated by ChatGPT versus Google in clinical scenarios involving a new diagnosis of GYN cancer.
Methods:
Clinical scenarios representing early- and advanced-stage endometrial, ovarian, and cervical cancers were developed by GYN oncologists using publicly available patient education materials. Each included four standardized questions on etiology, prognosis, treatment, and treatment efficacy. ChatGPT-4 and Google were queried for each question, with new sessions for ChatGPT and private browsing for Google to minimize bias. Responses were independently rated by four blinded GYN oncology experts. Accuracy was scored on a 6-point Likert scale; completeness and reference quality were scored on 3-point scales. Reference quality was categorized as low (commercial), medium (institutional/government), or high (peer-reviewed). Reviewer comments were also collected. Descriptive statistics and Wilcoxon signed-rank tests were used for analysis.
Results:
Across six clinical scenarios and 21 standardized questions (n=84 total responses), ChatGPT outperformed Google on all evaluated domains. The median accuracy score was 6 (IQR 5-6) for ChatGPT and 5 (IQR 4-6) for Google (P=.04). The median completeness score was 3 (IQR 3-3) for ChatGPT and 2 (IQR 1-3) for Google (P=.009). Reference quality was also higher for ChatGPT with a median score of 3 (IQR 3-3) compared to 2 (IQR 2-2) for Google (P=.009). Reviewer comments noted that ChatGPT provided more accurate, comprehensive, and personalized responses, while Google returned less detailed content from general consumer health websites rather than peer-reviewed sources.
Conclusions:
ChatGPT may serve as a reliable and high-quality resource for patient education in GYN oncology. Compared to Google, it delivered more accurate, complete, and better-referenced content. Further research is needed to assess the readability and accessibility of ChatGPT-generated content across diverse patient populations.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.