Previously submitted to: Journal of Medical Internet Research (no longer under consideration since May 04, 2023)
Date Submitted: May 3, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
How does ChatGPT4 preform on Non-English National Medical Licensing Examination? An Evaluation in Chinese Language
ABSTRACT
Background:
ChatGPT, an artificial intelligence (AI) system powered by large-scale language models, has garnered significant interest in the healthcare. Its performance dependent on the quality and amount of training data available for specific language. This study aims to assess the of ChatGPT's ability in medical education and clinical decision-making within the Chinese context.
Objective:
Evaluate ChatGPT's performance on the Chinese NMLE conducted within the Chinese context.
Methods:
We utilized a dataset from the Chinese National Medical Licensing Examination (NMLE) to assess ChatGPT-4's proficiency in medical knowledge within the Chinese language. Performance indicators, including score, accuracy, and concordance (confirmation of answers through explanation), were employed to evaluate ChatGPT's effectiveness in both original and encoded medical questions. Additionally, we translated the original Chinese questions into English to explore potential avenues for improvement.
Results:
ChatGPT scored 442/600 for original questions in Chinese, surpassing the passing threshold of 360/600. However, ChatGPT demonstrated reduced accuracy in addressing open-ended questions, with an overall accuracy rate of 47.7%. Despite this, ChatGPT displayed commendable consistency, achieving a 75% concordance rate across all case analysis questions. Moreover, translating Chinese case analysis questions into English yielded only marginal improvements in ChatGPT's performance (P =0.728).
Conclusions:
ChatGPT exhibits remarkable precision and reliability when handling the NMLE in Chinese language. Translation of NMLE questions from Chinese to English does not yield an improvement in ChatGPT's performance.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.