Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 8, 2026
Date Accepted: Sep 7, 2026

The final, peer-reviewed published version of this preprint can be found here:

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

Luo S, Tang R, Chen Z, Wang J, Hou H, Ma L, Liu M

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

J Med Internet Res 2026;28:e100710

DOI: 10.2196/100710

PMID: 42826371

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

In the Era of GPT Image 2, Lower Barriers to Conclusion-Oriented Falsification Challenge Integrity

  • Shuhang Luo; 
  • Runhua Tang; 
  • Ziyin Chen; 
  • Jianye Wang; 
  • Huimin Hou; 
  • Li Ma; 
  • Ming Liu

ABSTRACT

Background:

As generative artificial intelligence (GenAI) advances, generated images are becoming increasingly difficult to distinguish from authentic ones. This may be particularly concerning for conclusion-oriented falsified experimental images, which can be derived from genuine results and altered to support a desired hypothesis while remaining visually plausible.

Objective:

To evaluate whether the latest image generative AI, namely GPT image 2 can create conclusion-altered biomedical images from authentic published figures and to test whether human readers with certain research experience and commercial AI-detection tools could identify such conclusion-oriented falsified images.

Methods:

This diagnostic evaluation study assembled a benchmark dataset of 104 images, including 52 authentic images from 16 published biomedical studies and 52 matched AI-manipulated counterparts. Manipulated images were generated to preserve visual plausibility while reversing the apparent experimental direction. Image types included western blots and subcutaneous xenograft tumor photographs. Three commercial AI-detection tools were evaluated in parallel with 24 PhD-level researchers with more than 5 years of research experience, each of whom classified 50 randomly sampled images. AI-manipulated biomedical images generated from authentic published figures using a contemporary image-generation workflow.For human readers, outcomes were accuracy, sensitivity, and specificity for classifying images as authentic or AI manipulated. For automated tools, the main outcome was area under the receiver operating characteristic curve (AUC) with 95% CIs.

Results:

Among 24 human readers, mean accuracy was 50.5% (SD, 6.9%), with sensitivity of 38.5% (SD, 18.9%) and specificity of 61.5% (SD, 17.7%), indicating near-chance discrimination. Among automated tools, AI or Not showed the best overall performance (AUC, 0.790; 95% CI, 0.695-0.885), followed by Is It AI (AUC, 0.733; 95% CI, 0.645-0.821), whereas Illuminarty performed at chance level (AUC, 0.485; 95% CI, 0.372-0.598).

Conclusions:

In this study, generative AI produced visually credible biomedical images that altered apparent experimental conclusions and were not reliably recognized by expert readers. Current detection tools showed limited standalone utility, supporting stronger safeguards such as raw-data retention, provenance requirements, and source-file review.


 Citation

Please cite as:

Luo S, Tang R, Chen Z, Wang J, Hou H, Ma L, Liu M

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

J Med Internet Res 2026;28:e100710

DOI: 10.2196/100710

PMID: 42826371

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.