Currently submitted to: JMIR Research Protocols
Date Submitted: Aug 5, 2026
Open Peer Review Period: Aug 6, 2026 - Oct 1, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Comparing Large Language Models for Risk-of-Bias Assessment in Randomized Controlled Trials Using the Cochrane RoB 2 Tool: Protocol for a Reliability and Agreement Study
ABSTRACT
Background:
The Cochrane RoB 2 tool is the current standard for risk-of-bias assessment in randomized trials, but its application is time-consuming, shows low inter-rater reliability, and is frequently misapplied even under Cochrane editorial oversight. Large language models (LLMs) have shown preliminary promise for this task, yet existing evidence is limited by the absence of direct head-to-head comparisons between models, reliance on a single input format, and inadequate control for within-review clustering. Critically, no study has evaluated whether LLMs meet the performance thresholds required for specific operational roles in evidence synthesis, such as high-risk triage or low-risk confirmation, which demand diagnostic accuracy metrics rather than global agreement statistics.
Objective:
To compare two current LLMs, under two input modalities, against Cochrane authors' consensus-based RoB 2 assessments as the reference standard, and to determine whether they meet pre-specified thresholds for three distinct operational roles: adjunct reviewer, high-risk triage, and low-risk confirmation.
Methods:
This pre-registered reliability and agreement study evaluates four conditions: two LLMs (Claude Sonnet 4.6 and ChatGPT 5.4) crossed with two input modalities (full-text paste and PDF upload). A sampling design of one RCT per systematic review prevents within-review clustering. The target sample is 350 RCTs, powered to estimate weighted Cohen κ with a 95% CI half-width of ±0.10. The primary outcome is agreement of each LLM condition with the reference standard for the global RoB 2 judgment. Secondary outcomes include per-domain agreement (5 domains), inter-model and inter-modality agreement, intra-LLM test-retest stability (2 iterations), diagnostic performance of the dichotomized judgment (sensitivity, specificity, PPV, NPV, LR+, LR−), and time efficiency. Analyses use Python (≥3.10; scipy, numpy, pingouin, sklearn). Pre-specified operational thresholds are: adjunct reviewer, κ ≥ 0.40; high-risk triage, sensitivity ≥ 0.80 and LR+ ≥ 5.0; low-risk confirmation, specificity ≥ 0.90 and LR− ≤ 0.20. Prompt design is informed by documented Cochrane misapplication patterns (Moore et al., 2023). The literature search closure date is 31 May 2026.
Results:
The protocol was pre-registered on the Open Science Framework (DOI 10.17605/OSF.IO/QDHKJ), and an ethics exemption was granted by the Ethics Committee of the Hospital Clínico de la Universidad de Chile. Data collection is scheduled to begin in September 2026, with two evaluators assessing complementary RCT sets to avoid modality–evaluator confounding.
Conclusions:
This study will provide the first direct, head-to-head comparison of current LLMs and input modalities for RoB 2 assessment, benchmarked against role-specific operational thresholds rather than global agreement alone. The findings will clarify whether, and in which capacity, LLMs can be responsibly integrated into evidence synthesis workflows. Clinical Trial: Not a clinical trial. Pre-registered on the Open Science Framework: https://doi.org/10.17605/OSF.IO/QDHKJ
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.