Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Jul 14, 2023)

Date Submitted: Apr 30, 2023

Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.

Improving Detection of ChatGPT-Generated Fake BioMedical Science Using Real Publication Text: Introducing xFakeBibs a Supervised-Learning Network Algorithm

  • Ahmed Hamed; 
  • Xindong Wu

ABSTRACT

Background:

ChatGPT is becoming a new reality. Where do we go from here?

Objective:

is to show how we can distinguish ChatGPT-generated publications from counterparts produced by biomedical scientists.

Methods:

By means of a new algorithm, called xFakeBibs, we show the significant difference between ChatGPT-generated fake publications and real publications. Specifically, we triggered ChatGPT to generate 100 publications that were related to Alzheimer’s disease and comorbidity. Using the TF-IDF measure against a dataset of real publications, we constructed a network training model of the bigrams extracted from 100 publications. By 10-folds of 100 publications each, we built 10 calibrating networks to derive lower/upper bounds for classifying an article as real or fake. The final step of the algorithm is designed to test xFakeBibs against each of the ChatGPT-generated articles and predict its class. The xFakeBibs algorithm successfully assigned the POSITIVE label for real and NEGATIVE for fake ones.

Results:

When comparing the training model with the calibration models, we found that the similarities fluctuated between (19%-21%) of bigram overlaps. The calibrating folds contributed (51%-70%) of new bigrams, while ChatGPT contributed only 23% (> 50% of any of the other 10 calibrating folds). When classifying the individual articles, the xFakeBibs algorithm predicted 98/100 publications as fake, while 2 articles failed the test and were classified as real publications.

Conclusions:

This work provided clear evidence on how to distinguish ChatGPT-generated articles from real articles. The analysis demonstrated how such contents are distinguishable in bulk. Also, the algorithmic approach demonstrated the detection the individual fake articles with a high degree of accuracy. However, it remains challenging to detect all fake records. ChatGPT may seem to be a useful tool, but it certainly presents a threat to our authentic knowledge and real science. This work is indeed a step in the right direction to counter fake science and misinformation. Clinical Trial: N/A


 Citation

Please cite as:

Hamed A, Wu X

Improving Detection of ChatGPT-Generated Fake BioMedical Science Using Real Publication Text: Introducing xFakeBibs a Supervised-Learning Network Algorithm

JMIR Preprints. 30/04/2023:48604

DOI: 10.2196/preprints.48604

URL: https://preprints.jmir.org/preprint/48604

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.