Accepted for/Published in: Journal of Medical Internet Research
Date Submitted: Mar 27, 2026
Date Accepted: Aug 3, 2026
Structured Guidance Improves Clinical Practice Guideline Appraisal by Human Experts and AI Agents: A Systematic Review, Meta-Analysis, and Validation Study
ABSTRACT
Background:
The number of rehabilitation clinical practice guidelines (CPGs) has grown rapidly, yet their methodological rigor varies widely, which constrains implementation. Established appraisal tools such as AGREE II and RIGHT offer structured standards for evaluation, but their application is time-intensive and difficult to scale. Large language models (LLMs) - based AI agents may offer a scalable alternative, although their reliability in rehabilitation guideline appraisal remains unclear.
Objective:
This study aimed to assess the current methodological and reporting quality of global rehabilitation CPGs and examine whether an AI agent based on LLMs, supported by standardized training protocols, can achieve performance comparable to human experts.
Methods:
We conducted a systematic review of rehabilitation CPGs published in global databases up to May 2025. Methodological quality was evaluated using AGREE II, and reporting quality using the RIGHT checklist. Determinants of guideline quality were analyzed through univariate analyses, multivariable logistic regression to identify independent predictors, and subgroup analyses of reporting rates. To test the AI agent performance, we implemented a 2×2 factorial experiment comparing DeepSeek-R1 and O1-mini with human expert consensus under zero-shot and Workbook-supported conditions. External validation was performed using six anterior cruciate ligament reconstruction CPGs, with agreement against reference standards and appraisal time recorded.
Results:
We included 205 CPGs (147 English-language, 58 Chinese-language). After introducing a standardized Guideline Appraisal Manual, human AGREE II agreement improved markedly, with mean intraclass correlation coefficients (ICCs) increasing from (−0.09 – 0.66) to (0.82 – 0.91) across domains. Overall guideline quality was low. Applicability (35.1%) and Stakeholder Involvement (50.7%) were the weakest domains. English-language guidelines scored higher than Chinese-language guidelines in Scope and Purpose (72.3±14.3 vs. 66.4±12.0, P=0.002) and Applicability (37.8±17.5 vs. 28.3±14.7, P<0.001). External review was the strongest independent predictor of high quality (OR 2.69; 95% CI 2.18–3.32). RIGHT assessments showed consistent reliability (ICC 0.76–0.86). Reporting was highest for Basic Information (68.6%) and lowest for Funding, Declaration, and Management of Interests (42.7%). Guidelines published before release of the RIGHT checklist demonstrated lower overall reporting scores (logOR -0.15; 95% CI -0.28 to -0.01). Under zero-shot conditions, Agent–expert agreement was moderate (ICC approximately 0.61). Workbook integration improved performance, with DeepSeek-R1 increasing from 0.614 to 0.726 and surpassing O1-mini (0.679). In external validation, DeepSeek-R1 maintained stable agreement (ICC 0.711) and completed each appraisal in 5.44 minutes, compared with 11.18 minutes for human reviewers, without sacrificing accuracy.
Conclusions:
Rehabilitation CPGs show persistent methodological gaps, particularly in applicability and cross-regional consistency. Zero-shot LLM appraisal is insufficient. With structured training support, however, AI agents, particularly DeepSeek-R1, can provide reliable and efficient assistance in guideline evaluation, supporting scalable translation of evidence into practice. Clinical Trial: The protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO; registration number: CRD420251270676).
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.