Currently submitted to: JMIR Medical Education
Date Submitted: Aug 28, 2026
Open Peer Review Period: Aug 28, 2026 - Oct 23, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Transcript-based large language model scoring as a low-cost second rater for standardized patient history-taking examinations: a retrospective adjudication study
ABSTRACT
Background:
Objective scoring of standardized patient (SP) history-taking examinations typically relies on a single on-site examiner, whose real-time ratings are subject to leniency, fatigue, and halo effects. A second human rater would roughly double faculty workload. Large language models (LLMs) can score encounter transcripts at modest marginal cost.
Objective:
We evaluated a workflow in which examiner and LLM score independently, concordant items are auto-accepted, and discordant items are routed to expert review.
Methods:
We retrospectively analysed 240 student-SP encounter transcripts from a diagnostic history-taking final examination (six cases; 8,405 checklist items). Each item was scored by the on-site examiner and independently by an LLM (DeepSeek; temperature 0) using a predefined rubric. Scores agreed on 7,523 items (89.5%); the 882 discordant items were re-scored by three experts, blinded to both original scores, and majority consensus (≥2 of 3) served as the primary reference standard, with full-consensus and single-expert sensitivity analyses. A stratified 10% sample of concordant items (n = 752) was adjudicated to estimate the shared-error rate. Two further models replicated the scoring.
Results:
Experts reached a majority on 92.7% of discordant items (Fleiss κ = 0.45). Against majority consensus (n = 818), the LLM matched the reference on 65.5% of items (95% CI 62.2-68.7) versus 27.9% (24.9-31.0) for the examiner; the 37.7-percentage-point difference was robust under student-level cluster bootstrap (95% CI 30.7-44.3) and in all sensitivity analyses. Mean absolute error was 0.37 versus 0.75 points. Where experts awarded zero, the examiner had also awarded zero in only 26.1% of items versus 69.3% for the LLM, indicating examiner leniency. Findings replicated with a lightweight API model and an on-premises open-weight model (49.9% and 50.6%). Blinded expert review confirmed 99.1% of audited auto-accepted scores. After expert correction, 29.2% of examiner grade-band classifications changed versus 19.2% under LLM scoring.
Conclusions:
Auto-accepting concordant items and adjudicating only the discordant minority concentrates expert attention where single-rater scoring is most error-prone while reducing expert review workload by ~90%. Transcript-based LLM scoring is a feasible, low-cost second rater and quality-assurance complement; whether it can replace a second human rater requires reference standards independent of the transcript modality.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.