Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: JMIR Medical Education (no longer under consideration since May 21, 2024)

Date Submitted: Sep 3, 2023

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

A Comparative Study on the Effectiveness of three Generative Pre-trained Transformer Models in Assessing Spinal Metastasis from both Physician and Patient Perspectives

  • Wei Xu; 
  • Jianru Xiao; 
  • Xiang Wang; 
  • Bo Li; 
  • Guanyu Fang; 
  • Zhaoyu Chen; 
  • Jiefu Fan

ABSTRACT

Background:

With technological advancements, large language models (LLMs) like ChatGPT, NewBing, and Google Bard are emerging in medicine.

Objective:

This study examines their application in orthopedics, particularly spinal metastases diagnosis and treatment, assessing their efficacy in aiding surgeons and informing patients.

Methods:

The study utilized questions from both doctor and patient viewpoints. We derived 15 questions on spinal metastases surgical treatment from the doctor's perspective, encompassing preoperative diagnosis to postoperative care. Additionally, 30 patient concerns about spinal metastases were integrated, representing outpatient and inpatient uncertainties. These questions were posed to ChatGPT, NewBing, and Google Bard. Two orthopedic surgeons assessed each response, categorizing them using a five-point Likert scale.

Results:

ChatGPT addressed all 39 questions with 58.97% in Strong Agreement, 30.77% Agreement, and lesser percentages in other categories. NewBing's results were 25.64% Strong Agreement, 48.72% Agreement, and varied for others. Google Bard had 17.95% Strong Agreement and 53.85% Agreement. For doctor-centric questions, ChatGPT led in Strong Agreement at 58.9%, trailed by NewBing (25.6%) and Google Bard (17.9%). ChatGPT notably surpassed NewBing and Google Bard in comprehensive doctor responses (OR = 6.000, P = .006). For patient queries, the three LLMs showed comparable performance in Strong Agreement and Agreement.

Conclusions:

The research underscores ChatGPT's proficiency in supporting orthopedic decisions on spinal metastasis. All models exhibited parallel performance for patient queries. While LLMs provide essential perspectives, they should complement, not replace, medical expertise.


 Citation

Please cite as:

Xu W, Xiao J, Wang X, Li B, Fang G, Chen Z, Fan J

A Comparative Study on the Effectiveness of three Generative Pre-trained Transformer Models in Assessing Spinal Metastasis from both Physician and Patient Perspectives

JMIR Preprints. 03/09/2023:52409

DOI: 10.2196/preprints.52409

URL: https://preprints.jmir.org/preprint/52409

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.