Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently submitted to: JMIR AI

Date Submitted: Oct 1, 2026
Open Peer Review Period: Oct 7, 2026 - Dec 2, 2026
(currently open for review)

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Evaluating Expressed Clinical Reasoning in Large Language Model Outputs: A Viewpoint on a Proposed Assessment Framework

  • Zhangshu Joshua Jiang; 
  • Zina Ibrahim; 
  • James T Teo

ABSTRACT

Large language models (LLMs) are increasingly explored for clinical documentation, information synthesis, health communication, and decision-support tasks. Evaluation must address the quality of their outputs in clinically meaningful terms, not only whether a response reaches an expected answer. Endpoint correctness can conceal important failures in problem representation, differential prioritisation, response to new evidence, temporal synthesis, uncertainty, and safety. Existing clinical benchmarks address overlapping aspects of these qualities through physician-authored criteria and task-specific scoring, but organise them in different ways. Medical education offers established approaches to making selected aspects of clinical reasoning observable for assessment. Its instruments were developed for human learners: their validity does not transfer automatically to LLMs, and a generated explanation does not establish how a model produced its answer. We therefore use medical-education assessment traditions as sources of constructs and task-design logic, not as prevalidated scoring tools for model outputs. We propose a provisional seven-domain framework for evaluating expressed clinical reasoning: the clinically defensible organisation, interpretation, and updating of available patient information in a generated response. The domains cover accuracy and grounding, validity and completeness, diagnostic and management reasoning, temporal reasoning, uncertainty and updating, clinical safety, and communication. They are organised into an evidential foundation, an expressed-reasoning core, and clinical guardrails. Proposed five-point behavioural anchors are applied only to relevant domains and components. Case-specific reference criteria permit defensible alternatives, and judgements are made using evidence available at the specified decision point. Prespecified safety-critical breaches are flagged separately and cannot be compensated for by strong scores elsewhere. The framework is intended for offline evaluation, error analysis, and benchmark design. It is not a validated measurement instrument, a measure of internal model reasoning, or evidence of readiness for clinical deployment. We outline priorities for independent content review, feasibility testing, domain-level agreement assessment, and evaluation of automated judging before stronger measurement claims are made.


 Citation

Please cite as:

Jiang ZJ, Ibrahim Z, Teo JT

Evaluating Expressed Clinical Reasoning in Large Language Model Outputs: A Viewpoint on a Proposed Assessment Framework

JMIR Preprints. 01/10/2026:113478

DOI: 10.2196/preprints.113478

URL: https://preprints.jmir.org/preprint/113478

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.