Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Education

Date Submitted: Aug 30, 2026
Open Peer Review Period: Aug 31, 2026 - Oct 26, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Comparing AI-Driven Meta-Debriefing Against Clinical Human Experts in Simulation: A Proof-of-Concept Study

  • Kai-Yuan Cheng; 
  • Ibrahim Gomaa; 
  • Usha Asirvatham; 
  • Aidan F. Mullan; 
  • Diana Kelm; 
  • Daniel Cabrera

ABSTRACT

Background:

Meta-debriefing involves evaluating the debriefer with an emphasis on enhancing future debriefing activities. However, assessing debriefing effectiveness presents challenges stemming from both time limitations and resource demands.

Objective:

The aim of this project was to develop an Artificial Intelligence (AI)-driven meta-debriefer that parallels human meta-debriefers in the context of audio transcription through objective structured assessment of debriefing (OSAD) metrics.

Methods:

We conducted prospective observational cohort research at the Mayo Clinic Multidisciplinary Simulation Center in Rochester, gathering 29 simulation sessions after removing task trainer and simulation activities without debriefing sessions. The debriefing sessions were recorded and evaluated by an AI-driven meta-debriefer and two experienced human raters using OSAD metrics. The primary result was the non-inferiority of the OSAD score of AI compared to human evaluation.

Results:

This study utilized a predefined non-inferiority margin of 10% (-0.40 per domain, -2.80 total score). The AI model showed comparable score to human experts for the total score (mean difference: -2.22; 95% CI: -3.05 to -1.39, noninferiority p = 0.91) while significantly underscored the Approach and Engagement of learner domains (mean difference -0.86, 95% CI:-0.97 to -0.76, noninferiority p<0.001; mean difference -0.66, 95% CI: -0.87 to -0.44, noninferiority p = 0.010).

Conclusions:

This proof-of-concept study demonstrates the potential of LLM-based meta-debriefers. While current models overscore “Application” and underscore “Approach” and “Engagement of Learners” due to emotional nuance limitations, they represent a valuable supplementary tool, offering a roadmap to enhance clinical educators' efficiency.


 Citation

Please cite as:

Cheng KY, Gomaa I, Asirvatham U, Mullan AF, Kelm D, Cabrera D

Comparing AI-Driven Meta-Debriefing Against Clinical Human Experts in Simulation: A Proof-of-Concept Study

JMIR Preprints. 30/08/2026:93971

DOI: 10.2196/preprints.93971

URL: https://preprints.jmir.org/preprint/93971

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.