Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Sep 29, 2025
Date Accepted: Aug 12, 2026
Evaluating the Accuracy of Large Language Models in Risk of Bias Assessment Using the ROB 2.0 Tool: Exploratory Feasibility Study
ABSTRACT
Background:
Large language models (LLMs) have the potential to improve the efficiency of evidence synthesis, but their reliability in performing complex tasks such as risk of bias (ROB) assessment in randomized controlled trials (RCTs) remains unclear.
Objective:
This study aimed to evaluate whether LLMs can reliably assess ROBs in RCTs using the ROB 2.0 tool.
Methods:
This survey study was conducted between December 28, 2024, and February 28, 2025, in adherence with American Association for Public Opinion Research (AAPOR) reporting guidelines. 29 RCTs were selected from published Cochrane systematic reviews (SRs) across diverse medical fields. Each RCT was independently evaluated twice by ChatGPT, with Cochrane review authors’ assessments serving as the reference standard for comparison. The main outcomes included the accuracy and consistency of ROB 2.0 assessments at both domain and study levels, including accuracy metrics such as correct assessment rate, sensitivity, specificity, and F1 score. Consistency between the repeated assessments was quantified using Cohen’s κ and prevalence-adjusted bias-adjusted κ (PABAκ).
Results:
The LLM demonstrated a moderate overall assessment rate of 75.2% (95% CI, 67.4%-83.0%) in the first assessment and 75.9% (95% CI, 66.6%-85.2%) in the second. Sensitivity decreased from 66.8% to 54.3%, while specificity increased from 77.6% to 81.2%. Domain-level correct assessment rate ranged from 63.8% to 87.9%, with the lowest accuracy and F1 scores observed in domain 1. Overall consistency between repeated assessments was high, with a mean agreement of 90% (SD, 0.08) and Cohen κ values were 0.93, 0.84, and 0.85 in Domain 1, 4, and 5.
Conclusions:
In this survey, ChatGPT demonstrated moderate accuracy and high consistency in assessing ROB in RCTs using the ROB 2.0 tool. These findings suggest that LLMs may support methodological evaluations in systematic reviews.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.