Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: Journal of Medical Internet Research

Date Submitted: Jul 19, 2026
Open Peer Review Period: Jul 22, 2026 - Sep 16, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Enhancing Clinical Training in Oral and Maxillofacial Surgery Using an LLM-Based Intelligent Standardized Patient System: A Randomized Controlled Trial

  • Jiayu Shen; 
  • Yao Yuan; 
  • Sichen Han; 
  • Bingxin Fan; 
  • Zilin Wang; 
  • Xinliang Duan

ABSTRACT

Background:

Large language model-based intelligent standardized patient systems offer a scalable solution for clinical skills training, yet rigorous comparisons against active pedagogical controls remain scarce, and the mechanisms underlying their effectiveness are poorly specified. This study evaluated an LLM-based intelligent SP system compared with small-group, tutor-facilitated case discussion in dental education, with the explicit objective of disentangling the role of differential individual active learning time as a potential mediator of observed effects.

Objective:

To evaluate an LLM-based intelligent standardized patient system compared with small-group, tutor-facilitated case discussion in dental education, and to disentangle the role of differential individual active learning time as a potential mediator of observed effects.

Methods:

In this single-center, parallel, open-label randomized trial conducted from May to June 2025, 60 third- and fourth-year dental students were randomized 1:1 to the AI-SP group (n=30) or the TC-SG group (n=30). The AI-SP group completed six interactive virtual patient modules powered by DeepSeek-V3 with automated feedback, while the TC-SG group discussed identical cases in groups of four to five students with a tutor. The primary outcome was the adjusted post-intervention mini-CEX Overall Score assessed by a blinded expert panel using ANCOVA with baseline adjustment. Secondary outcomes included AI-generated communication metrics and clinical self-efficacy. The study design inherently produced substantially different individual active learning time between arms—approximately 110–120 minutes per session for AI-SP versus 24–30 minutes for TC-SG—which we explicitly treated as a design-defined dose parameter. Due to significant baseline imbalances in secondary AI-generated metrics favouring the TC-SG group, causal inferences were restricted to the primary outcome, where baseline balance was confirmed.

Results:

Fifty-nine participants completed the trial (AI-SP=30, TC-SG=29). For the primary outcome, the AI-SP group showed significantly higher adjusted post-intervention mini-CEX Overall Score compared with TC-SG (adjusted mean difference [aMD]=0.68; 95% CI, 0.29–1.07; P=0.001; η²p=0.22). This effect size corresponded to the substantial difference in individual practice density between the two conditions. For AI-generated secondary metrics, despite significant within-group improvements in the AI-SP group (all P<0.001), no significant between-group differences were observed at post-intervention for Accuracy (P=0.059) or Interactivity (P=0.161), indicating comparable endpoint performance. Given the baseline imbalance in these metrics, we interpret these null between-group differences conservatively, without claims of catch-up or superiority. An exploratory regression analysis treating estimated individual active learning time as a continuous predictor revealed a significant association with mini-CEX improvement (β=0.34, P=0.002), supporting the dose-response interpretation.

Conclusions:

The AI-SP system, by affording substantially higher individual active learning density, produced superior expert-assessed clinical outcomes compared with small-group discussion. However, this advantage is confounded by unequal practice time and cannot be attributed solely to AI intelligence. Our findings suggest that AI-SP functions primarily as an effective pedagogical platform for scaling individual deliberate practice opportunities. Future research must employ dose-equated designs to isolate the unique contribution of AI-generated feedback from the general benefits of increased active learning time.


 Citation

Please cite as:

Shen J, Yuan Y, Han S, Fan B, Wang Z, Duan X

Enhancing Clinical Training in Oral and Maxillofacial Surgery Using an LLM-Based Intelligent Standardized Patient System: A Randomized Controlled Trial

JMIR Preprints. 19/07/2026:107462

DOI: 10.2196/preprints.107462

URL: https://preprints.jmir.org/preprint/107462

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.