Currently submitted to: JMIR Formative Research
Date Submitted: Sep 27, 2026
Open Peer Review Period: Sep 27, 2026 - Nov 22, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Teaching Critical Appraisal of AI-Generated Output in Medical Education: A Mixed Methods Pilot Study
ABSTRACT
Background:
Medical students need practical opportunities to critically appraise artificial intelligence (AI)-generated clinical output. Comparing model interpretations with an independently formed judgment offers a way to teach this skill within clinical education.
Objective:
To examine learner acceptability, explore changes in confidence and automation-bias-related self-reports, and identify useful features and refinements of an electrocardiogram (ECG)-based AI-appraisal session.
Methods:
This single-center mixed methods pilot involved 13 medical and biomedical sciences students at a US osteopathic medical school. Groups interpreted three ECGs independently, consulted Claude and ChatGPT using a standardized prompt, questioned and compared their outputs, and discussed reference interpretations with faculty. Paired surveys assessed confidence, automation-bias-related responses, learner evaluations, recalled behavior, mental effort, and future intentions. Exact two-sided Wilcoxon signed-rank tests with Holm adjustment across 10 exploratory tests were used for seven confidence and three wording-matched awareness-related items. Two differently worded awareness items and confidence composites were summarized descriptively. Free-text responses underwent AI-assisted descriptive qualitative content analysis and were integrated narratively with survey findings.
Results:
All 13 participants provided complete paired survey data; 11 contributed 40 free-text responses. All participants endorsed curricular inclusion and found the live interactions engaging; 12/13 (92.3%) would recommend the session. The largest mean confidence increase concerned justifying agreement or disagreement with AI, from 4.31 to 6.38 on a 0-10 scale (change 2.08; unadjusted 95% CI 0.77-3.62; Holm-adjusted P=.08). Error-recognition confidence increased from 3.85 to 5.15 (adjusted P=.42). The descriptive AI-appraisal confidence summary increased by 1.69 points, while the ECG-reading summary changed by −0.02 points. No paired test met the multiplicity-adjusted .05 threshold. Six qualitative categories highlighted comparison, questioning and verification, independent reasoning and self-monitoring, preparation, improvements, and collaboration. Six of 11 participants named model comparison as the most useful element, and 11 of 13 endorsed its value for refocusing participants on their own reasoning.
Conclusions:
This pilot identifies a well-received approach to teaching critical appraisal of AI-generated output. Learners valued model comparison and reported opportunities for independent reasoning and reflection, with the largest descriptive confidence changes in AI appraisal. The findings provide a practical instructional sequence and learner-informed refinements for curriculum development and subsequent evaluation of performance and transfer. Clinical Trial: Not applicable.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.