Accepted for/Published in: JMIR Research Protocols
Date Submitted: Jun 1, 2026
Date Accepted: Aug 6, 2026
Process-Oriented, Behaviourally Anchored Assessment of Clinical Reasoning in Large Language Models and the Effect of Extended Thinking (REACT-AI): Protocol for a Prospective, Multigroup Comparative Study
ABSTRACT
Background:
Most evaluations of clinical reasoning in large language models (LLMs) score only the final answer, usually the accuracy of multiple-choice questions. This provides little information about how the model reasons, which is important for safety. Two recent developments have made this gap pressing: reasoning-optimised models are now common, and several systems expose explicit, extended thinking control. Whether this extra deliberation actually improves the reasoning process, and metacognition in particular, has not been tested using a validat-ed instrument.
Objective:
This protocol aims to achieve three objectives. This study introduces and validates REACT-AI, a behaviourally anchored rating scale (BARS) with 13 sub-domains for the process-oriented assessment of AI clinical reasoning. It builds a conflict-of-interest-controlled, cross-tier LLM-as-Judge pipeline for scoring at scale. It also tests whether extended thinking improves rea-soning quality.
Methods:
This study was a prospective and comparative design with two phases and a within-model thinking-mode factor. Six flagship models were run in standard and extended thinking or reasoning modes, resulting in 12 conditions. Four providers are toggled inside the same model (Anthropic Claude Opus 4.8, Google Gemini 3.1 Pro, Alibaba Qwen 3.7 Max, xAI Grok 4.3). Two providers contrasted a standard model with a dedicated reasoning model (OpenAI GPT-5.5 versus o3; DeepSeek V4 Pro versus R1). In Phase 1, five standardised urgent-care vi-gnettes were answered under all 12 conditions, three runs each, for 180 AI outputs. The same vignettes were answered by an expert clinician panel (at least 10), senior medical stu-dents (30 to 71), and junior students (30 to 100). Every output was scored twice: by dual-blind trained human raters and, in parallel, by LLM judges. Phase 1 calibrates the automated scores against human experts, with targets of Pearson r of at least 0.80, weighted kappa of at least 0.60, and ICC above 0.75. In Phase 2, the validated judge pipeline scored 3,600 AI out-puts (100 vignettes by 12 conditions by three runs). No model judges its own family, and judges sit one tier below the candidates, which blocks self-enhancement bias. A parallel shadow track measures the Self-Preference Bias Delta. The secondary instruments were the AI Metacognition Assessment Rubric (AI-MAR) and the Disinformation Generation Rate (DGR). Gulf English prompts probe the robustness of sociolinguistic variation.
Results:
This is a study protocol. As of submission, institutional ethics approval has been granted (United Arab Emirates University Social Sciences Ethics Committee, ERSC_2025_6124), and the study has been preregistered on the Open Science Framework, although no data have yet been collected. Data collection is expected to begin in [month, year], with Phase 1 cali-bration anticipated by June 2026 and the Phase 2 benchmark by August 2026. The study will report the validated REACT-AI instrument, the first multi-model BARS benchmark of clinical reasoning, calibration data on how closely LLM judges track human experts, self-preference bias estimates for each model family, and the process-level effect of extended thinking on the Reflection and Metacognition domain, including any null findings.
Conclusions:
REACT-AI and its judge pipeline are built to be reusable. They close the gap between accura-cy benchmarks and genuine reasoning assessments and provide the biomedical informatics community with a template for studying reasoning mode effects. Clinical Trial: Open Science Framework
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.