Previously submitted to: JMIR Medical Education (no longer under consideration since May 21, 2024)
Date Submitted: Sep 3, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
A Comparative Study on the Effectiveness of three Generative Pre-trained Transformer Models in Assessing Spinal Metastasis from both Physician and Patient Perspectives
ABSTRACT
Background:
With technological advancements, large language models (LLMs) like ChatGPT, NewBing, and Google Bard are emerging in medicine.
Objective:
This study examines their application in orthopedics, particularly spinal metastases diagnosis and treatment, assessing their efficacy in aiding surgeons and informing patients.
Methods:
The study utilized questions from both doctor and patient viewpoints. We derived 15 questions on spinal metastases surgical treatment from the doctor's perspective, encompassing preoperative diagnosis to postoperative care. Additionally, 30 patient concerns about spinal metastases were integrated, representing outpatient and inpatient uncertainties. These questions were posed to ChatGPT, NewBing, and Google Bard. Two orthopedic surgeons assessed each response, categorizing them using a five-point Likert scale.
Results:
ChatGPT addressed all 39 questions with 58.97% in Strong Agreement, 30.77% Agreement, and lesser percentages in other categories. NewBing's results were 25.64% Strong Agreement, 48.72% Agreement, and varied for others. Google Bard had 17.95% Strong Agreement and 53.85% Agreement. For doctor-centric questions, ChatGPT led in Strong Agreement at 58.9%, trailed by NewBing (25.6%) and Google Bard (17.9%). ChatGPT notably surpassed NewBing and Google Bard in comprehensive doctor responses (OR = 6.000, P = .006). For patient queries, the three LLMs showed comparable performance in Strong Agreement and Agreement.
Conclusions:
The research underscores ChatGPT's proficiency in supporting orthopedic decisions on spinal metastasis. All models exhibited parallel performance for patient queries. While LLMs provide essential perspectives, they should complement, not replace, medical expertise.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.