Accepted for/Published in: JMIR Medical Education
Date Submitted: Feb 2, 2026
Date Accepted: Aug 11, 2026
Effect of Large Language Model–Powered Virtual Standardized Patients on History-Taking Among Undergraduate Medical Students: A Propensity-Matched Cohort Study
ABSTRACT
Background:
Medical history-taking (MHT) is a foundational competency for medical students, yet traditional standardized patient (SP)-based training faces challenges including high costs, management difficulties, and inconsistent standardization. Large language model-powered virtual standardized patients (LLM-VSPs) offer a potential solution by enabling scalable, standardized practice environments.
Objective:
This study aims to explore the effectiveness of using LLM-VSPs in enhancing medical students' MHT skills and the correlation between their practice behaviors and the improvement of MHT skills.
Methods:
A cohort of 168 third-year medical students was allocated to an LLM-VSPs intervention group (Group A, n=120) or a control group (Group B, n=48). Group A completed MHT training using an LLM-VSPs system with real-time feedback, while Group B followed standard curricula. Propensity score matching (PSM) balanced baseline characteristics between groups. Outcomes were assessed through standardized MHT scoring (total score 100: 60 for content, 40 for communication skills). Practice data in Group A were analyzed for correlations with final performance.
Results:
Post-PSM analysis (n=40 per group) demonstrated balanced baselines (standardized mean differences <0.1). Group A achieved significantly higher total MHT scores than Group B (87.7 ± 7.29 vs. 83.7 ± 8.06, p = .023), particularly in terms of the content of MHT (50.2 ± 5.65 vs. 47.4 ± 5.93, p = .031). Within Group A, total practice duration (r = 0.183, p = .046) and average AI-generated assessment score(r = 0.233, p = .010) positively correlated with the final scores, while practice frequency showed no significant association.
Conclusions:
This study demonstrates that LLM-VSPs can enhance MHT skill development more effectively than traditional methods, primarily through enabling deliberate, feedback-driven practice rather than repetitive task accumulation.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.