Accepted for/Published in: JMIR Medical Education
Date Submitted: Apr 18, 2026
Date Accepted: Aug 24, 2026
AI-Assisted Angoff Standard Setting for Multiple-Choice Examinations in Medical Education: A Comparative Study of Faculty Judges and Large Language Models
ABSTRACT
Background:
Standard setting is essential for a defensible assessment in medical education. The modified Angoff method requires several expert judges and determining the “minimally competent candidate (MCC)” is cognitively challenging. However, empirical evidence for the role of artificial intelligence in standard setting is unclear. Therefore, we aimed to evaluate the use of large language models (LLMs) in Angoff standard setting compared with faculty judges.
Objective:
This study aims to examine the role of Large Language Models (LLMs) in the modified Angoff standard setting method for multiple-choice question (MCQ) in a basic medical sciences examination.
Methods:
This study was conducted in year 2 of the MD program in the United Arab Emirates. Ten faculty judges and five LLMs (ChatGPT, GROK, DeepSeek, MedGemma, and Claude Sonnet 4.5) determined the Angoff cut scores for a summative examination (120 MCQs). Standardized prompts were used for LLMs to mimic the same information provided for faculty judges.
Results:
Faculty-generated Angoff estimates were comparable to LLM-generated estimates, with no significant difference between groups (t (238) = .23, p = .818). In addition, the percent variance related to test items was much higher in LLMs compared with faculty judges (46.9% vs 18.4%). LLMs differed in their MCC conceptualization and approaches to determining item-level percentages of correct responses. However, rater-related variance was six times lower in LLMs compared with faculty judges (2.6% vs 16.4%). In addition, root mean square error (RMSE) of Angoff estimates was much lower in LLMs compared with faculty judges (1.28 vs 2.3). The correlations between Angoff estimates and item-related P-values were larger in LLMs compared with faculty judges (.552 vs .437). Furthermore, there were comparable pass rates by applying the cut scores generated by LLMs and faculty judges.
Conclusions:
Angoff estimates generated by LLMs were comparable to those of faculty judges, produced lower rater-related variance, and demonstrated greater precision in the estimated passing score. In addition, applying cut scores generated by LLM led to comparable pass rates with faculty with more sensitivity to item difficulty.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.