Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Currently accepted at: JMIR AI

Date Submitted: Feb 10, 2026
Open Peer Review Period: Feb 10, 2026 - Feb 11, 2026
Date Accepted: Jul 9, 2026
(closed for review but you can still tweet)

This paper has been accepted and is currently in production.

It will appear shortly on 10.2196/93239

The final accepted version (not copyedited yet) is in this tab.

Augmenting Oncology Guideline Maintenance with Large Language Models: A Prospective Case Study

  • Manuel Knauer; 
  • Julian Greß; 
  • Jakob Nikolas Kather; 
  • Peter May

ABSTRACT

Background:

Maintenance of oncology clinical practice guidelines (CPGs) is increasingly challenged by the rapid growth of trial data and therapeutic complexity. While large language models (LLMs) have shown promise in information retrieval, their utility in the rigorous, end-to-end workflow of guideline maintenance remains under-explored.

Objective:

This study aimed to systematically evaluate the performance of frontier LLMs in supporting oncology guideline maintenance. We sought to determine their reliability in predicting necessary guideline updates based on new evidence, their accuracy in extracting data from clinical trials, and their effectiveness as automated auditors for detecting errors in established guidelines.

Methods:

Using the Onkopedia peripheral T-cell lymphoma (PTCL) guideline as a natural experiment, we tasked frontier models with deep-research modes (Gemini 2.5 Pro, GPT o4-mini-high) to predict a guideline update in August 2025 based on the 2021 version. Predictions were validated against the official 2025 revision published in October 2025. Next, we benchmarked evidence extraction accuracy across 80 pivotal trials using models of varying scale (27B–671B parameters vs. frontier). Finally, we deployed a stacked LLM workflow to audit 28 recently updated Onkopedia guidelines for linguistic and content-related errors.

Results:

In the predictive task, models captured 36.7–40% of substantive updates, often identifying landmark approvals but frequently overstating evidence. While frontier models demonstrated high accuracy (up to 99.2%) in extracting data from individual studies, substantially outperforming smaller open-source models, this precision declined during multi-source synthesis. As automated auditors of existing CPGs, the models successfully identified a median of 16.5 formal errors per document and detected several clinically relevant inconsistencies (e.g., invalid scoring formulas, incorrect staging definitions).

Conclusions:

LLMs currently lack the reasoning stability for autonomous guideline authoring due to deficits in complex synthesis. However, they are effective tools for high-fidelity evidence extraction and automated quality assurance, supporting a human-led, AI-augmented workflow for efficient guideline maintenance.


 Citation

Please cite as:

Knauer M, Greß J, Kather JN, May P

Augmenting Oncology Guideline Maintenance with Large Language Models: A Prospective Case Study

JMIR AI. 09/07/2026:93239 (forthcoming/in press)

DOI: 10.2196/93239

URL: https://preprints.jmir.org/preprint/93239

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.