Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR AI

Date Submitted: Mar 17, 2026
Date Accepted: Jul 3, 2026

The final, peer-reviewed published version of this preprint can be found here:

Evidence Use and Identifier-Conditioned Prior Knowledge in Large Language Model Classification of Oncology Trials Assessed Through Progressive Content Removal and Counterfactual Testing: Comparative Analysis

Windisch P, Koechli C, Dennstädt F, Aebersold DM, Zwahlen DR, Förster R, Schröder C

Evidence Use and Identifier-Conditioned Prior Knowledge in Large Language Model Classification of Oncology Trials Assessed Through Progressive Content Removal and Counterfactual Testing: Comparative Analysis

JMIR AI 2026;5:e95565

DOI: 10.2196/95565

PMID: 42525868

PMCID: 13419288

Do Large Language Models Use Provided Evidence or Identifier-Conditioned Prior Knowledge? Progressive Content Removal and Counterfactual Testing in Oncology Trial Classification: A Comparative Analysis

  • Paul Windisch; 
  • Carole Koechli; 
  • Fabio Dennstädt; 
  • Daniel M. Aebersold; 
  • Daniel R. Zwahlen; 
  • Robert Förster; 
  • Christina Schröder

ABSTRACT

Background:

Large language models (LLMs) can classify biomedical documents accurately, but strong performance does not prove they are using the supplied text rather than identifier-triggered parametric knowledge.

Objective:

To test whether oncology trial-success classification reflects “reading” of abstract evidence or “remembering” of known trials.

Methods:

We used a corpus of 250 two-arm oncology randomized controlled trials from seven major journals (2005 - 2023) and asked the flagship models of three commercial vendors (OpenAI, Google, and Anthropic) to output a single label indicating whether the primary endpoint was met. For each trial we created five deterministic inputs: title+abstract (baseline), title-only, DOI-only, counterfactual title+abstract with the primary endpoint outcome minimally flipped, and the same counterfactual title+abstract paired with the original DOI to induce an identifier-text conflict.

Results:

With full title+abstract, models achieved near-ceiling performance (accuracy and F1 Score 0.96 - 0.97) and high format adherence (97.2 - 100%). Performance degraded stepwise with content removal (title-only accuracy and F1 Score 0.79 - 0.88, DOI-only 0.63 - 0.67), consistent with above-chance identifier-driven signal. Under counterfactual results, models followed the edited evidence (accuracy and F1 Score 0.96 - 0.99 against inverted labels). Adding the real DOI minimally affected GPT (accuracy and F1 Score ≈ 0.99) but modestly reduced Gemini (accuracy and F1 Score ≈ 0.97) and Claude (accuracy and F1 Score ≈ 0.95), mainly via lower sensitivity.

Conclusions:

LLMs robustly track explicit endpoint statements in abstracts, yet identifiers can support above-chance predictions and occasionally compete with textual evidence. Progressive ablations plus counterfactual conflicts provide a practical, reproducible audit for grounding in biomedical LLM evaluations.


 Citation

Please cite as:

Windisch P, Koechli C, Dennstädt F, Aebersold DM, Zwahlen DR, Förster R, Schröder C

Evidence Use and Identifier-Conditioned Prior Knowledge in Large Language Model Classification of Oncology Trials Assessed Through Progressive Content Removal and Counterfactual Testing: Comparative Analysis

JMIR AI 2026;5:e95565

DOI: 10.2196/95565

PMID: 42525868

PMCID: 13419288

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.