Currently submitted to: Journal of Medical Internet Research
Date Submitted: Jul 19, 2026
Open Peer Review Period: Jul 22, 2026 - Sep 16, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Enhancing Clinical Training in Oral and Maxillofacial Surgery Using an LLM-Based Intelligent Standardized Patient System: A Randomized Controlled Trial
ABSTRACT
Background:
Large language model-based intelligent standardized patient systems offer a scalable solution for clinical skills training, yet rigorous comparisons against active pedagogical controls remain scarce, and the mechanisms underlying their effectiveness are poorly specified. This study evaluated an LLM-based intelligent SP system compared with small-group, tutor-facilitated case discussion in dental education, with the explicit objective of disentangling the role of differential individual active learning time as a potential mediator of observed effects.
Objective:
To evaluate an LLM-based intelligent standardized patient system compared with small-group, tutor-facilitated case discussion in dental education, and to disentangle the role of differential individual active learning time as a potential mediator of observed effects.
Methods:
In this single-center, parallel, open-label randomized trial conducted from May to June 2025, 60 third- and fourth-year dental students were randomized 1:1 to the AI-SP group (n=30) or the TC-SG group (n=30). The AI-SP group completed six interactive virtual patient modules powered by DeepSeek-V3 with automated feedback, while the TC-SG group discussed identical cases in groups of four to five students with a tutor. The primary outcome was the adjusted post-intervention mini-CEX Overall Score assessed by a blinded expert panel using ANCOVA with baseline adjustment. Secondary outcomes included AI-generated communication metrics and clinical self-efficacy. The study design inherently produced substantially different individual active learning time between arms—approximately 110–120 minutes per session for AI-SP versus 24–30 minutes for TC-SG—which we explicitly treated as a design-defined dose parameter. Due to significant baseline imbalances in secondary AI-generated metrics favouring the TC-SG group, causal inferences were restricted to the primary outcome, where baseline balance was confirmed.
Results:
Fifty-nine participants completed the trial (AI-SP=30, TC-SG=29). For the primary outcome, the AI-SP group showed significantly higher adjusted post-intervention mini-CEX Overall Score compared with TC-SG (adjusted mean difference [aMD]=0.68; 95% CI, 0.29–1.07; P=0.001; η²p=0.22). This effect size corresponded to the substantial difference in individual practice density between the two conditions. For AI-generated secondary metrics, despite significant within-group improvements in the AI-SP group (all P<0.001), no significant between-group differences were observed at post-intervention for Accuracy (P=0.059) or Interactivity (P=0.161), indicating comparable endpoint performance. Given the baseline imbalance in these metrics, we interpret these null between-group differences conservatively, without claims of catch-up or superiority. An exploratory regression analysis treating estimated individual active learning time as a continuous predictor revealed a significant association with mini-CEX improvement (β=0.34, P=0.002), supporting the dose-response interpretation.
Conclusions:
The AI-SP system, by affording substantially higher individual active learning density, produced superior expert-assessed clinical outcomes compared with small-group discussion. However, this advantage is confounded by unequal practice time and cannot be attributed solely to AI intelligence. Our findings suggest that AI-SP functions primarily as an effective pedagogical platform for scaling individual deliberate practice opportunities. Future research must employ dose-equated designs to isolate the unique contribution of AI-generated feedback from the general benefits of increased active learning time.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.