Accepted for/Published in: JMIR Medical Education
Date Submitted: Dec 17, 2024
Date Accepted: May 15, 2026
Can AI tell right from wrong? A comparison of GPT-4o and Claude’s performance in ethics scenarios
ABSTRACT
Background:
The emergence of ChatGPT has sparked curiosity regarding the capabilities of AI in the field of medicine. Minimal research exists regarding the proficiency of ChatGPT in ethical scenarios, specifically in specialty-based scenarios. This study aimed to compare the performance of ChatGPT on ethics questions with that of medical students and orthopaedic residents.
Objective:
This study aims to assess the ability of ChatGPT to answer ethical and legal scenario questions at the level of medical students and orthopaedic residents.
Methods:
A total of 201 ethical/legal scenario questions were randomly generated from question banks targeted for third and fourth-year medical students (UWorld, AMBOSS), and orthopaedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into ChatGPT 4o three separate times. If answers varied between trials, the answer provided most frequently by ChatGPT was used as the selected answer.
Results:
ChatGPT provided different answers to identically worded trials for 20.33% of general scenario questions and 4.31% of orthopaedic-specific questions. ChatGPT scored an average of 69% and 71.75% on general ethics and orthopaedic–specific questions respectively, with no significant difference compared to the performance of medical students or orthopaedic residents (p=.70). If ChatGPT selected the incorrect response, it chose the incorrect response most commonly chosen by medical students 62.05% of the time, and by orthopaedic residents 72.02% of the time.
Conclusions:
These results indicate that ChatGPT can answer both general and specialty-specific ethical and legal questions with proficiency on par with both medical students and orthopaedic residents. Our findings also suggest that ChatGPT may err in interpretation of these concepts and questions in a similar manner to human learners. However, the observed variance and contradictory outputs in response to identical inputs is concerning and raises the need for further investigation. These findings carry significant implications for the validity of all AI-related research and should be considered when designing future studies.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.