Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: Journal of Medical Internet Research

Date Submitted: May 8, 2026
Date Accepted: Sep 7, 2026

The final, peer-reviewed published version of this preprint can be found here:

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

Luo S, Tang R, Chen Z, Wang J, Liu M, Wang J, Ma L

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

J Med Internet Res 2026;28:e100710

DOI: 10.2196/100710

PMID: 42826371

Conclusion-Altered Biomedical Research Images Generated by AI: A Diagnostic Study of Human and Automated Detection

  • Shuhang Luo; 
  • Runhua Tang; 
  • Ziyin Chen; 
  • Jianye Wang; 
  • Ming Liu; 
  • Jianfeng Wang; 
  • Li Ma

ABSTRACT

Background:

Generative artificial intelligence (GenAI) is increasingly integrated into scientific workflows. While offering significant utility, its capacity to fabricate convincing images in specific biomedical experimental settings raises serious concerns for research integrity. Specifically, the ability to seamlessly alter experimental trends to support false conclusions poses a critical, yet under-evaluated, threat.

Objective:

This study aimed to assess whether a contemporary high-fidelity generative model (ChatGPT Images 2.0) could generate “conclusion-altered” images from authentic biomedical figures in two selected modalities-western blot and subcutaneous xenograft tumor images-that preserve visual characteristics while reversing apparent experimental outcomes, and to determine whether such images could be reliably identified by uncalibrated human experts or commercial AI-detection tools.

Methods:

We constructed a benchmark dataset of 104 images, comprising 52 authentic images from 16 published biomedical studies and 52 matched AI-manipulated counterparts in western blot and subcutaneous xenograft tumor subgroups. Authentic images were iteratively refined using natural-language prompts via “ChatGPT Images 2.0” without post-processing. The diagnostic performance of three commercial AI-detection tools (“Illuminarty”, “Is It AI”, and “AI or Not”) was evaluated using receiver operating characteristic (ROC) analyses. In parallel, 24 PhD-level biomedical researchers (active peer reviewers familiar with these image types) independently classified images under routine conditions without prior forensic training.

Results:

In the evaluated settings, human readers showed near-chance performance in distinguishing authentic from AI-manipulated images, with mean accuracy of 50.5% (SD 6.9%), sensitivity of 38.5% (SD 18.9%), specificity of 61.5% (SD 17.7%), and mean Youden index of 0. Among automated tools, “AI or Not” achieved the best discrimination (AUC=0.790; 95% CI 0.695-0.885), followed by “Is It AI” (AUC=0.733; 95% CI 0.645-0.821); “Illuminarty” performed at chance level (AUC=0.485; 95% CI 0.372-0.598). Although some tools showed moderate discrimination, performance remained insufficient for reliable application in these settings, given false-positive risks and unclear thresholds.

Conclusions:

In the evaluated western blot and xenograft settings, current high-fidelity GenAI can produce conclusion-altered images that evade routine human expert scrutiny and challenge tested detection tools. These findings highlight vulnerabilities in specific biomedical image modalities and underscore the need for enhanced safeguards. We propose a three-step editorial workflow for these settings: mandating original uncropped raw images at submission, screening metadata during triage, and instructing reviewers to focus on biological inconsistencies.


 Citation

Please cite as:

Luo S, Tang R, Chen Z, Wang J, Liu M, Wang J, Ma L

Biomedical Research Images Manipulated by Generative AI to Alter Scientific Outcomes: Diagnostic Study of Human and Automated Detection

J Med Internet Res 2026;28:e100710

DOI: 10.2196/100710

PMID: 42826371

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.