Currently submitted to: JMIR AI
Date Submitted: Oct 1, 2026
Open Peer Review Period: Oct 7, 2026 - Dec 2, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Evaluating Expressed Clinical Reasoning in Large Language Model Outputs: A Viewpoint on a Proposed Assessment Framework
ABSTRACT
Large language models (LLMs) are increasingly explored for clinical documentation, information synthesis, health communication, and decision-support tasks. Evaluation must address the quality of their outputs in clinically meaningful terms, not only whether a response reaches an expected answer. Endpoint correctness can conceal important failures in problem representation, differential prioritisation, response to new evidence, temporal synthesis, uncertainty, and safety. Existing clinical benchmarks address overlapping aspects of these qualities through physician-authored criteria and task-specific scoring, but organise them in different ways. Medical education offers established approaches to making selected aspects of clinical reasoning observable for assessment. Its instruments were developed for human learners: their validity does not transfer automatically to LLMs, and a generated explanation does not establish how a model produced its answer. We therefore use medical-education assessment traditions as sources of constructs and task-design logic, not as prevalidated scoring tools for model outputs. We propose a provisional seven-domain framework for evaluating expressed clinical reasoning: the clinically defensible organisation, interpretation, and updating of available patient information in a generated response. The domains cover accuracy and grounding, validity and completeness, diagnostic and management reasoning, temporal reasoning, uncertainty and updating, clinical safety, and communication. They are organised into an evidential foundation, an expressed-reasoning core, and clinical guardrails. Proposed five-point behavioural anchors are applied only to relevant domains and components. Case-specific reference criteria permit defensible alternatives, and judgements are made using evidence available at the specified decision point. Prespecified safety-critical breaches are flagged separately and cannot be compensated for by strong scores elsewhere. The framework is intended for offline evaluation, error analysis, and benchmark design. It is not a validated measurement instrument, a measure of internal model reasoning, or evidence of readiness for clinical deployment. We outline priorities for independent content review, feasibility testing, domain-level agreement assessment, and evaluation of automated judging before stronger measurement claims are made.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.