Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Medical Education (no longer under consideration since Mar 08, 2025)

Date Submitted: Nov 12, 2024

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Performance of GPT-4 and mainstream Chinese Large Language Models on the Chinese Postgraduate Examination dataset: Potential for AI-assisted Traditional Chinese Medicine

  • Suyuan Peng; 
  • Yan Zhu; 
  • Baifeng Wang; 
  • Meiwei Zhang; 
  • Zhe Wang; 
  • Keyu Yao; 
  • Meng Hao; 
  • Junhui Wang

ABSTRACT

Background:

In China, the medical education system is characterized by multiple co-existing levels. Physicians with higher levels of education typically have better job prospects. Consequently, the medical master’s degree examination holds greater significance in the selection process compared to the Chinese licensing examination. The application of Large Language Models(LLMs) in Traditional Chinese Medicine (TCM) has rapidly expanded and intensified. Theories of TCM carry distinct scientific significance, requiring LLMs to have advanced information processing and comprehension abilities within a Chinese language context. LLMs have performed notably well in the medical licensing exams of many countries. However, their performance in selective examinations within TCM still requires further investigation.

Objective:

The study aimed to comprehensively evaluate and compare the performance of Ernie Bot, ChatGLM, SparkDesk, and GPT-4 in processing the 2023 Chinese Postgraduate Examination for Traditional Chinese Medicine (TCM) questions, and to explore their potential applications in the TCM field.

Methods:

The performance of the 4 mainstream Large Language Models(LLMs), namely Ernie Bot, ChatGLM, SparkDesk, and GPT-4, were evaluated using the 2023 Chinese Postgraduate Examination questions for TCM as a test dataset. We calculated the exam scores, displaying LLM's performance on various subjects, and evaluated the output responses based on three qualitative metrics: logical reasoning, and the ability to use internal and external information.

Results:

Ernie Bot and ChatGLM both achieved accuracy rates of 50.30% and 46.67%, respectively, which were over the passing score. There was a statistically significant difference in performance across test subjects observed among Ernie Bot, ChatGLM, and GPT-4, with the highest performance achieved in the module of medical humanistic spirit. Logical reasoning: ChatGLM and GPT-4 provided logical explanations in each response for the answer selection, whereas Ernie Bot and SparkDesk showed logical reasoning in 98.2% and 43.6% of responses, respectively. Internal information: Both ChatGLM and GPT-4 incorporated internal information into all answer explanations, whereas SparkDesk exhibited a notably low quantity of responses that used internal information. External information: Over 60% of responses from Ernie Bot, ChatGLM, and GPT-4 included external information. The utilization of external information did not significantly differ between correct and incorrect answers. Additionally, there was a statistically significant difference in the percentage of correct answers in SparkDesk based on the presence of internal or external information (P< .001).

Conclusions:

Ernie Bot and ChatGLM's expertise in TCM surpassed the passing threshold for the postgraduate selection examination. The impressive ability of LLMs to reason logically and integrate background information demonstrated the substantial potential of LLMs in TCM.


 Citation

Please cite as:

Peng S, Zhu Y, Wang B, Zhang M, Wang Z, Yao K, Hao M, Wang J

Performance of GPT-4 and mainstream Chinese Large Language Models on the Chinese Postgraduate Examination dataset: Potential for AI-assisted Traditional Chinese Medicine

JMIR Preprints. 12/11/2024:68152

DOI: 10.2196/preprints.68152

URL: https://preprints.jmir.org/preprint/68152

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.