Accepted for/Published in: JMIR Medical Education
Date Submitted: Mar 15, 2026
Date Accepted: Sep 8, 2026
Large Language Model Performance on Multistep Clinical Cases: Comparative Study Across Question and Case Levels
ABSTRACT
Background:
Most Large Language Models (LLMs) have achieved passing scores on medical licensing examinations. However, most evaluations predominantly focus on single-question accuracy, overlooking the necessity of consistency across multi-step patient management scenarios, such as making a diagnosis followed by a treatment plan. It is unclear if LLMs can maintain logical consistency across these multi-step cases.
Objective:
This study aims to evaluate the consistency of LLMs on multi-step clinical reasoning case, and to investigate the impact of model size scaling and case complexity on performance stability.
Methods:
We curated a dataset of 189 unique clinical cases comprising 473 individual questions from the Chinese National Medical Licensing Examination. The number of questions related to each case ranges from 2 to 4. We tested four representative LLMs, including DeepSeek-R1, GPT-4o, Gemini-2.5-Flash, and Qwen2.5. We measured two metrics: the Question Pass Rate (QPR) and the Case Pass Rate (CPR). A case was considered correct only if the model answered all its questions correctly. The consistency gap was calculated as the difference between QPR and CPR. Additionally, we tested Qwen2.5 at different parameter sizes (3B to 32B) to see if model size affects stability.
Results:
All LLMs achieved high scores on individual questions with QPR exceeding 83%. However, the CPR were significantly lower. DeepSeek-R1 demonstrated the highest stability with a QPR of 89.85% and a CPR of 79.89%. This performance corresponded to the smallest consistency gap of 9.96%. In contrast, GPT-4o exhibited a QPR of 83.51% and a CPR of 65.61%, resulting in the largest consistency gap of 17.9%. Regarding model scaling, we found that increasing the model size from 3B to 32B reduced the consistency gap by approximately 50%. Furthermore, as cases became longer and more complex, the consistency gap increased for all LLMs.
Conclusions:
Current LLMs exhibit high accuracy in answering single questions, but this does not mean LLMs can handle multi-step complex reasoning processes involved in clinical cases. Reasoning consistency is heavily dependent on model scale and case complexity. Therefore, LLMs should be used as support tools rather than independent decision-makers in medical education and clinical practice. Clinical Trial: None
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.