Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Education

Date Submitted: Mar 19, 2026
Date Accepted: Jul 30, 2026
Date Submitted to PubMed: Jul 31, 2026

The final, peer-reviewed published version of this preprint can be found here:

Evaluating Large Language Model–Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study

Gao X, Huang X, Hu R, Zhang L, Liu H, Zhang B, Wei C, Qiu W, Zhang M, Sun X

Evaluating Large Language Model–Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study

JMIR Med Educ 2026;12:e95578

DOI: 10.2196/95578

PMID: 42535722

Evaluating Large Language Model–Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: A Cross-Sectional Agreement Study

  • Xiaoxing Gao; 
  • Xiaoming Huang; 
  • Rongrong Hu; 
  • Li Zhang; 
  • Huiting Liu; 
  • Bingqing Zhang; 
  • Chong Wei; 
  • Wei Qiu; 
  • Mengyu Zhang; 
  • Xuefeng Sun

ABSTRACT

Background:

Large language model–powered virtual standardized patients enable scalable clinical skills practice, but the validity of AI-generated performance scores—and whether they align sufficiently with faculty ratings across different raters and cases—remains unclear for determining appropriate educational use.

Objective:

To evaluate agreement between large language model (LLM)–generated and faculty ratings of history-taking and communication performance in an LLM-powered virtual standardized patient (VSP) platform, and to examine how rater and case heterogeneity influence agreement to define appropriate use boundaries for LLM-based scoring.

Methods:

In this cross-sectional observational study, 92 fourth-year medical students completed a 15-minute voice-based VSP encounter across three cases, and 10 faculty raters scored performance using a standardized rubric (total 0–100; information gathering and communication each 0–50). We evaluated AI–faculty associations, agreement, and rater effects using linear mixed-effects models (random rater intercepts), along with intraclass correlation coefficients (ICC), Spearman correlations, mean absolute error (MAE), mixed-effects Bland–Altman analyses, and variance partition coefficients (VPC).

Results:

Median total scores were similar for AI and faculty (93.0 [IQR 6.0] vs 94.0 [IQR 4.0]), with substantial rater variability in faculty total scores (VPC=0.37). AI total scores were positively associated with faculty total scores (β=0.37, P<.001; Spearman ρ=0.50, 95% CI 0.34–0.65). Absolute agreement for total score was moderate (ICC[2,1]=0.51, 95% CI 0.34–0.65), with MAE 3.11 points. Mixed-effects Bland–Altman analysis showed an adjusted mean bias (Human−AI) of 1.26 points and 95% limits of agreement from −4.95 to 7.48, with proportional bias (β_mean=−0.55, P<.001). Agreement was stronger for information gathering than communication (information gathering: β=0.46, ρ=0.49, ICC=0.54, VPC=0.23; communication: β=0.27, ρ=0.28, ICC=0.29, VPC=0.52).

Conclusions:

LLM-based scoring in a VSP demonstrated moderate agreement with faculty ratings, performing better for information gathering than for communication. Given rater/case heterogeneity and proportional bias, this approach appears appropriate for formative feedback but not unsupervised high-stakes summative assessment. Further work should improve communication scoring and reduce score-dependent discrepancies; LLM scoring may also support faculty rater calibration.


 Citation

Please cite as:

Gao X, Huang X, Hu R, Zhang L, Liu H, Zhang B, Wei C, Qiu W, Zhang M, Sun X

Evaluating Large Language Model–Based Automated Scoring in a Voice-Based Virtual Standardized Patient Platform for Medical Students: Cross-Sectional Agreement Study

JMIR Med Educ 2026;12:e95578

DOI: 10.2196/95578

PMID: 42535722

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.