Previously submitted to: JMIR Medical Education (no longer under consideration since Aug 21, 2025)
Date Submitted: Feb 17, 2025
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Utilizing Large Language Models for Educating Patients About Polycystic Ovary Syndrome in China: A Two-Phase Study
ABSTRACT
Background:
Polycystic ovary syndrome (PCOS) is a prevalent condition requiring effective patient education, particularly in China. Large language models (LLMs) present a promising avenue for this. This two-phase study evaluates six LLMs for educating Chinese patients about PCOS. It assesses their capabilities in answering questions, interpreting ultrasound images, and providing patient instructions within a real-world clinical setting in China.
Objective:
This study systematically evaluated six gigantic language models—Gemini 2.0 Pro, OpenAI o1, ChatGPT-4o, ChatGPT-4, ERINE 4.0, and GLM-4—for use in gynecological medicine. It assessed their performance in several areas: answering questions from the Chinese Gynecology Qualification Examination, understanding gynecological ultrasound images, coping with polycystic ovary syndrome (PCOS) cases, writing patient instructions, and helping to solve clinical problems.
Methods:
A two-step evaluation method was used. Primarily, they tested the six frameworks on 136 gynecological exam questions and 36 ultrasound images. They then compared the results with those of medical students and residents. Six gynecologists evaluated the framework's responses to 23 PCOS-related questions using a Likert scale, and a Chinese readability tool was used to review the content objectively. In the following process, 40 patients with PCOS tested the two central systems, Gemini 2.0 Pro and OpenAI o1. They evaluated them in terms of patient satisfaction, text readability, and professional evaluation.
Results:
During the initial phase of testing, OpenAI o1 and Gemini 2.0 Pro demonstrated impressive accuracy on gynecological specialist exam questions, achieving rates of 93.63% and 92.40%, respectively. Additionally, their performance in ultrasound image diagnostic tasks was noteworthy, with OpenAI o1 achieving an accuracy of 69.44% and Gemini 2.0 Pro reaching 53.70%. Regarding the response to PCOS-related questions, OpenAI o1 significantly outperformed the other models in accuracy, completeness, readability, practicality, and safety. However, its responses were notably more complex (average score 13.98, p = 0.003). The second-phase evaluation revealed that Gemini 2.0 Pro excelled in readability (patient rating 3.45, p < 0.01; physician rating 3.35, p = 0.03), significantly surpassing OpenAI o1 (patient rating 2.65, physician rating 2.90). However, Gemini 2.0 Pro slightly lagged behind OpenAI o1 in response completeness (3.05 vs. 3.50, p = 0.04).
Conclusions:
This study reveals that large language models (LLMs) have considerable potential to address the issues faced by Chinese patients with PCOS, particularly OpenAI o1 and Gemini 2.0 Pro, which are capable of providing accurate and comprehensive responses. Nevertheless, it still needs to be strengthened so that it can balance clarity and comprehensiveness. In addition, the performance of other big language models besides OpenAI o1 in analyzing ultrasound images, especially the ability to handle regulation categories, needs to be improved to meet the regulation of medical practice. Clinical Trial: None
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.