Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Education

Date Submitted: Dec 17, 2024
Date Accepted: May 15, 2026

The final, peer-reviewed published version of this preprint can be found here:

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study

Desai KR, Gorsky AL, Zelenski NA

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study

JMIR Med Educ 2026;12:e70199

DOI: 10.2196/70199

PMID: 42492056

PMCID: 13396000

Can AI tell right from wrong? A comparison of GPT-4o and Claude’s performance in ethics scenarios

  • Karishma R. Desai; 
  • Anna L. Gorsky; 
  • Nicole A. Zelenski

ABSTRACT

Background:

The emergence of ChatGPT has sparked curiosity regarding the capabilities of AI in the field of medicine. Minimal research exists regarding the proficiency of ChatGPT in ethical scenarios, specifically in specialty-based scenarios. This study aimed to compare the performance of ChatGPT on ethics questions with that of medical students and orthopaedic residents.

Objective:

This study aims to assess the ability of ChatGPT to answer ethical and legal scenario questions at the level of medical students and orthopaedic residents.

Methods:

A total of 201 ethical/legal scenario questions were randomly generated from question banks targeted for third and fourth-year medical students (UWorld, AMBOSS), and orthopaedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into ChatGPT 4o three separate times. If answers varied between trials, the answer provided most frequently by ChatGPT was used as the selected answer.

Results:

ChatGPT provided different answers to identically worded trials for 20.33% of general scenario questions and 4.31% of orthopaedic-specific questions. ChatGPT scored an average of 69% and 71.75% on general ethics and orthopaedic–specific questions respectively, with no significant difference compared to the performance of medical students or orthopaedic residents (p=.70). If ChatGPT selected the incorrect response, it chose the incorrect response most commonly chosen by medical students 62.05% of the time, and by orthopaedic residents 72.02% of the time.

Conclusions:

These results indicate that ChatGPT can answer both general and specialty-specific ethical and legal questions with proficiency on par with both medical students and orthopaedic residents. Our findings also suggest that ChatGPT may err in interpretation of these concepts and questions in a similar manner to human learners. However, the observed variance and contradictory outputs in response to identical inputs is concerning and raises the need for further investigation. These findings carry significant implications for the validity of all AI-related research and should be considered when designing future studies.


 Citation

Please cite as:

Desai KR, Gorsky AL, Zelenski NA

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study

JMIR Med Educ 2026;12:e70199

DOI: 10.2196/70199

PMID: 42492056

PMCID: 13396000

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.