Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR Medical Education

Date Submitted: Aug 28, 2026
Open Peer Review Period: Aug 28, 2026 - Oct 23, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Transcript-based large language model scoring as a low-cost second rater for standardized patient history-taking examinations: a retrospective adjudication study

  • Jun Feng; 
  • Shaoting Wang; 
  • Xiaoxing Gao; 
  • Luo Wang; 
  • Xiaoming Huang; 
  • Xuefeng Sun

ABSTRACT

Background:

Objective scoring of standardized patient (SP) history-taking examinations typically relies on a single on-site examiner, whose real-time ratings are subject to leniency, fatigue, and halo effects. A second human rater would roughly double faculty workload. Large language models (LLMs) can score encounter transcripts at modest marginal cost.

Objective:

We evaluated a workflow in which examiner and LLM score independently, concordant items are auto-accepted, and discordant items are routed to expert review.

Methods:

We retrospectively analysed 240 student-SP encounter transcripts from a diagnostic history-taking final examination (six cases; 8,405 checklist items). Each item was scored by the on-site examiner and independently by an LLM (DeepSeek; temperature 0) using a predefined rubric. Scores agreed on 7,523 items (89.5%); the 882 discordant items were re-scored by three experts, blinded to both original scores, and majority consensus (≥2 of 3) served as the primary reference standard, with full-consensus and single-expert sensitivity analyses. A stratified 10% sample of concordant items (n = 752) was adjudicated to estimate the shared-error rate. Two further models replicated the scoring.

Results:

Experts reached a majority on 92.7% of discordant items (Fleiss κ = 0.45). Against majority consensus (n = 818), the LLM matched the reference on 65.5% of items (95% CI 62.2-68.7) versus 27.9% (24.9-31.0) for the examiner; the 37.7-percentage-point difference was robust under student-level cluster bootstrap (95% CI 30.7-44.3) and in all sensitivity analyses. Mean absolute error was 0.37 versus 0.75 points. Where experts awarded zero, the examiner had also awarded zero in only 26.1% of items versus 69.3% for the LLM, indicating examiner leniency. Findings replicated with a lightweight API model and an on-premises open-weight model (49.9% and 50.6%). Blinded expert review confirmed 99.1% of audited auto-accepted scores. After expert correction, 29.2% of examiner grade-band classifications changed versus 19.2% under LLM scoring.

Conclusions:

Auto-accepting concordant items and adjudicating only the discordant minority concentrates expert attention where single-rater scoring is most error-prone while reducing expert review workload by ~90%. Transcript-based LLM scoring is a feasible, low-cost second rater and quality-assurance complement; whether it can replace a second human rater requires reference standards independent of the transcript modality.


 Citation

Please cite as:

Feng J, Wang S, Gao X, Wang L, Huang X, Sun X

Transcript-based large language model scoring as a low-cost second rater for standardized patient history-taking examinations: a retrospective adjudication study

JMIR Preprints. 28/08/2026:110683

DOI: 10.2196/preprints.110683

URL: https://preprints.jmir.org/preprint/110683

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.