Previously submitted to: Journal of Medical Internet Research (no longer under consideration since Jul 14, 2023)
Date Submitted: Apr 30, 2023
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Improving Detection of ChatGPT-Generated Fake BioMedical Science Using Real Publication Text: Introducing xFakeBibs a Supervised-Learning Network Algorithm
ABSTRACT
Background:
ChatGPT is becoming a new reality. Where do we go from here?
Objective:
is to show how we can distinguish ChatGPT-generated publications from counterparts produced by biomedical scientists.
Methods:
By means of a new algorithm, called xFakeBibs, we show the significant difference between ChatGPT-generated fake publications and real publications. Specifically, we triggered ChatGPT to generate 100 publications that were related to Alzheimer’s disease and comorbidity. Using the TF-IDF measure against a dataset of real publications, we constructed a network training model of the bigrams extracted from 100 publications. By 10-folds of 100 publications each, we built 10 calibrating networks to derive lower/upper bounds for classifying an article as real or fake. The final step of the algorithm is designed to test xFakeBibs against each of the ChatGPT-generated articles and predict its class. The xFakeBibs algorithm successfully assigned the POSITIVE label for real and NEGATIVE for fake ones.
Results:
When comparing the training model with the calibration models, we found that the similarities fluctuated between (19%-21%) of bigram overlaps. The calibrating folds contributed (51%-70%) of new bigrams, while ChatGPT contributed only 23% (> 50% of any of the other 10 calibrating folds). When classifying the individual articles, the xFakeBibs algorithm predicted 98/100 publications as fake, while 2 articles failed the test and were classified as real publications.
Conclusions:
This work provided clear evidence on how to distinguish ChatGPT-generated articles from real articles. The analysis demonstrated how such contents are distinguishable in bulk. Also, the algorithmic approach demonstrated the detection the individual fake articles with a high degree of accuracy. However, it remains challenging to detect all fake records. ChatGPT may seem to be a useful tool, but it certainly presents a threat to our authentic knowledge and real science. This work is indeed a step in the right direction to counter fake science and misinformation. Clinical Trial: N/A
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.