Currently submitted to: JMIR Medical Education
Date Submitted: Sep 3, 2026
Open Peer Review Period: Sep 8, 2026 - Nov 3, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Error Direction of Large Language Models on Medical Multiple-Response Questions: An Opportunity-Standardized, Multimodel, Bilingual Benchmark Study
ABSTRACT
Background:
Large language models (LLMs) attain high accuracy on medical examinations, yet evaluations rely predominantly on single-best-answer multiple-choice questions (MCQs) and report only overall accuracy. Multiple-response items, termed X-type in China, require selecting every correct option and remain understudied. Whether models err by omitting correct options or by commissioning incorrect ones has yet to be assessed within an opportunity-standardized framework.
Objective:
To quantify the accuracy penalty of multiple-response (X-type) items relative to single-select items; to determine, after standardizing for the unequal numbers of correct and incorrect options, whether LLMs lean toward omission or commission; and to test whether option-level majority voting improves multiple-response performance.
Methods:
We analyzed 1,644 text-based items from ten examination cohorts (2017–2026) of the Chinese National Postgraduate Entrance Examination in Clinical Medicine, comprising 1,344 single-select items of types A1, A2, A3, and B, and 300 X-type items. Five LLMs, DeepSeek V4 Pro, Kimi K3, Doubao Seed-2.1-pro, GPT-5.6 Sol, and Gemini-3.6 Flash, were tested zero-shot bilingually, yielding 16,440 planned responses. Item-type effects were estimated using generalized estimating equations (GEE) with cluster-robust standard errors, adjusting for model, language, discipline, and cohort. Error direction was analyzed two ways: an option-level model treating each option as one judgment opportunity and the omission-bias index (OBI), defined as the within-response difference between proportions of correct options omitted and incorrect options selected.
Results:
Of 16,440 planned responses, 16,426 were evaluable. X-type accuracy (84.23%) was markedly below single-select accuracy (96.75%; adjusted OR 0.17, 95% CI 0.11–0.26; P<.0001), an adjusted gap of 13.6 percentage points, and the penalty was consistent across models and languages (item type × model and item type × language interactions, P=0.19 and P=0.82). Raw counts suggested more strict omissions (8.17%) than over-selections (6.70%), but X-type items averaged 2.79 correct versus 1.21 incorrect options per item, a 2.3-fold imbalance in error opportunities. After opportunity standardization, the per-opportunity false-selection rate was 6.69% (242/3,620), roughly twice the 3.57% per-opportunity omission rate (299/8,380; adjusted OR 0.38, 95% CI 0.20–0.70; P=0.002); pooled OBI was −0.039 (95% CI −0.062 to −0.018), concordant across conditions. This reversal showed no between-model difference (all Holm-adjusted pairwise comparisons, P≥.46). Option-level majority voting on the Chinese responses, selecting an option when at least 3 of 5 models chose it, reached 89.00% accuracy, above the best single model (87.00%), with an oracle upper bound of 96.00%.
Conclusions:
LLMs incur a substantial accuracy penalty on medical multiple-response items. The apparent omission dominance in raw counts is an artifact of answer-key opportunity imbalance; after opportunity standardization, models instead lean toward commission, a reversal consistent across models and languages. Option-level majority voting modestly outperformed the best single model. Medical LLM benchmarks should report accuracy stratified by item type and error direction on an opportunity-standardized basis.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.