Currently submitted to: JMIR Medical Education
Date Submitted: Aug 30, 2026
Open Peer Review Period: Aug 31, 2026 - Oct 26, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Comparing AI-Driven Meta-Debriefing Against Clinical Human Experts in Simulation: A Proof-of-Concept Study
ABSTRACT
Background:
Meta-debriefing involves evaluating the debriefer with an emphasis on enhancing future debriefing activities. However, assessing debriefing effectiveness presents challenges stemming from both time limitations and resource demands.
Objective:
The aim of this project was to develop an Artificial Intelligence (AI)-driven meta-debriefer that parallels human meta-debriefers in the context of audio transcription through objective structured assessment of debriefing (OSAD) metrics.
Methods:
We conducted prospective observational cohort research at the Mayo Clinic Multidisciplinary Simulation Center in Rochester, gathering 29 simulation sessions after removing task trainer and simulation activities without debriefing sessions. The debriefing sessions were recorded and evaluated by an AI-driven meta-debriefer and two experienced human raters using OSAD metrics. The primary result was the non-inferiority of the OSAD score of AI compared to human evaluation.
Results:
This study utilized a predefined non-inferiority margin of 10% (-0.40 per domain, -2.80 total score). The AI model showed comparable score to human experts for the total score (mean difference: -2.22; 95% CI: -3.05 to -1.39, noninferiority p = 0.91) while significantly underscored the Approach and Engagement of learner domains (mean difference -0.86, 95% CI:-0.97 to -0.76, noninferiority p<0.001; mean difference -0.66, 95% CI: -0.87 to -0.44, noninferiority p = 0.010).
Conclusions:
This proof-of-concept study demonstrates the potential of LLM-based meta-debriefers. While current models overscore “Application” and underscore “Approach” and “Engagement of Learners” due to emotional nuance limitations, they represent a valuable supplementary tool, offering a roadmap to enhance clinical educators' efficiency.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.