Currently submitted to: Journal of Medical Internet Research
Date Submitted: Jul 28, 2026
Open Peer Review Period: Jul 29, 2026 - Sep 23, 2026
(currently open for review)
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
The Systematic Assessment of GPT 5 versus Expert reviewers (SAGE) study: comparing human and zero-to-few shot GPT 5 Pro performance for every major systematic review and meta-analysis step.
ABSTRACT
Background:
Systematic reviews and meta-analyses (SRMAs) represent the highest bodies of evidence in clinical research, but the process is labor-intensive, averaging an estimated $141,194.80 and 11 months in the U.S.A. Large language models (LLMs) have emerged as tools to assist, and potentially automate, SRMA production. Reported LLM performance varies by task, model, and methodological design. Many studies evaluate only many-shot, task-specific performance requiring prompt engineering inaccessible to most researchers, limiting scalability. Conversely, fully agentic SRMA generation has proven unreliable, underscoring the need to identify performance bottlenecks requiring human intervention while automating the remaining tasks.
Objective:
To evaluate the performance of the GPT-5 Pro series across all major SRMA steps using a zero-to-few-shot pragmatic design, ensuring reproducibility for non-AI specialists.
Methods:
Full data for the search string, article screening, data extraction, risk of bias (ROB) assessment and data synthesis steps of 5 in-house SRMAs were extracted and processed. From October 2025 to February 2026, we prompted the GPT 5 Pro series through the ChatGPT interface, to perform each aforementioned SRMA step, given the original author data from the previous step. All prompts were designed using simple, open-sourced templates. For search string, article screening, and data synthesis steps, GPT outputs were compared to original author data. For data extraction and ROB assessment, GPT outputs were compared to a novel GPT-human hybrid reference standard. For each step, a review of human and LLM performance literature was undertaken to better contextualize results.
Results:
GPT’s search string sensitivity in capturing relevant references to be included in the final SRMAs was 85% [80-88] CI95%, landing within the range of reported human performances. GPT’s abstract screening sensitivity and specificity were respectively 80% [76-84] CI95% and 85% [84-86] CI95%, roughly comparable to humans. However, GPT’s full-text screening underperformed in sensitivity compared to humans, 46%, [41-52] CI95%, while maintaining a similar specificity, 99% [99-99] CI95%. GPT outperformed humans in data extraction and ROB assessment accuracy, respectively, 93% [90-95] CI95% and 83% [79-86] CI95%. Data synthesis’ meta-analytical results were meaningfully similar between GPT and original authors, but not always strictly identical. All tasks were completed 10 to 100 times faster by GPT than by its human counterparts.
Conclusions:
This study has limitations inherent to its zero-shot design, which often underperforms in highly “prompt-dependent” tasks, such as search string generation and article screening. Furthermore, in the absence of a better reference standard for evaluating GPT in certain steps, outputs were compared to original author data, which cannot be considered the gold-standard. In a zero-to-few shot pragmatic design, the GPT 5 Pro series performs comparably to humans in search string generation, abstract screening and data synthesis, while underperforming in full-text screening and outperforming humans in data extraction and ROB assessment. Clinical Trial: Clinical trial number: not applicable.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.