Accepted for/Published in: JMIR Medical Education
Date Submitted: Mar 19, 2026
Date Accepted: Jul 30, 2026
Date Submitted to PubMed: Jul 31, 2026
Evaluating Large Language Model–Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study
ABSTRACT
Background:
Large language model–powered virtual standardized patients enable scalable clinical skills practice, but the validity of AI-generated performance scores—and whether they align sufficiently with faculty ratings across different raters and cases—remains unclear for determining appropriate educational use.
Objective:
To evaluate agreement between large language model (LLM)–generated and faculty ratings of history-taking and communication performance in an LLM-powered virtual standardized patient (VSP) platform, and to examine how rater and case heterogeneity influence agreement to define appropriate use boundaries for LLM-based scoring.
Methods:
In this cross-sectional observational study, 92 fourth-year medical students completed a 15-minute voice-based VSP encounter across three cases, and 10 faculty raters scored performance using a standardized rubric (total 0–100; information gathering and communication each 0–50). We evaluated AI–faculty associations, agreement, and rater effects using linear mixed-effects models (random rater intercepts), along with intraclass correlation coefficients (ICC), Spearman correlations, mean absolute error (MAE), mixed-effects Bland–Altman analyses, and variance partition coefficients (VPC).
Results:
Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]), with substantial rater variability in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, P<.001; Spearman ρ=0.50, 95% CI 0.34–0.65). Absolute agreement for total score was moderate (ICC[2,1]=0.51, 95% CI 0.34–0.65), with MAE 3.11 points. Mixed-effects Bland–Altman analysis showed an adjusted mean bias (Human−AI) of 1.26 points and 95% limits of agreement from −4.95 to 7.48, with proportional bias (β_mean=−0.55, P<.001). Agreement was stronger for information gathering than communication (information gathering: β=0.46, ρ=0.49, ICC=0.54, VPC=0.23; communication: β=0.27, ρ=0.28, ICC=0.29, VPC=0.52).
Conclusions:
LLM-based scoring in a VSP demonstrated moderate agreement with faculty ratings, performing better for information gathering than for communication. Given rater/case heterogeneity and proportional bias, this approach appears appropriate for formative feedback but not unsupervised high-stakes summative assessment. Further work should improve communication scoring and reduce score-dependent discrepancies; LLM scoring may also support faculty rater calibration.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.