Previously submitted to: JMIR AI (no longer under consideration since Sep 03, 2023)
Date Submitted: Jun 17, 2023
Open Peer Review Period: Jun 17, 2023 - Aug 12, 2023
(closed for review but you can still tweet)
NOTE: This is an unreviewed Preprint
Warning: This is a unreviewed preprint (What is a preprint?). Readers are warned that the document has not been peer-reviewed by expert/patient reviewers or an academic editor, may contain misleading claims, and is likely to undergo changes before final publication, if accepted, or may have been rejected/withdrawn (a note "no longer under consideration" will appear above).
Peer review me: Readers with interest and expertise are encouraged to sign up as peer-reviewer, if the paper is within an open peer-review period (in this case, a "Peer Review Me" button to sign up as reviewer is displayed above). All preprints currently open for review are listed here. Outside of the formal open peer-review period we encourage you to tweet about the preprint.
Citation: Please cite this preprint only for review purposes or for grant applications and CVs (if you are the author).
Final version: If our system detects a final peer-reviewed "version of record" (VoR) published in any journal, a link to that VoR will appear below. Readers are then encourage to cite the VoR instead of this preprint.
Settings: If you are the author, you can login and change the preprint display settings, but the preprint URL/DOI is supposed to be stable and citable, so it should not be removed once posted.
Submit: To post your own preprint, simply submit to any JMIR journal, and choose the appropriate settings to expose your submitted version as preprint.
Warning: This is an author submission that is not peer-reviewed or edited. Preprints - unless they show as "accepted" - should not be relied on to guide clinical practice or health-related behavior and should not be reported in news media as established information.
Unmasking Assumptions- Evaluating the Impact of Data Augmentation on Machine Learning Models in Healthcare Datasets
ABSTRACT
Background:
Data imbalance is a critical issue in big data analysis, particularly when dealing with datasets containing fewer labels, such as healthcare real-world data, spam detection labels, and financial fraud detection datasets. While numerous data balancing methods have been proposed to enhance machine learning algorithm performance, research claims that Synthetic Minority Over-sampling Technique (SMOTE) and SMOTE-based data augmentation methods can improve algorithm performance. However, we observed that many online tutorials evaluating these methods use synthesized datasets, introducing bias into the evaluation process and leading to false positive improvements in performance.
Objective:
In this study, we propose a new evaluation framework for imbalanced data learning methods, experimenting with five data balancing techniques to assess their impact on machine learning algorithm performance.
Methods:
We collected 8 imbalanced real-world healthcare datasets with varying imbalance rates from different domains. We applied 6 data augmentation methods in conjunction with 11 machine learning techniques to test the efficacy of data augmentation in improving machine learning performance. Our proposed Evaluation Framework for Imbalanced Data Learning (EFIDL) uses a 5-fold cross-validation approach, comparing the traditional data augmentation evaluation methods with our new framework.
Results:
Traditional data augmentation evaluation methods can give a false impression of improved machine learning performance. However, our proposed evaluation framework demonstrates that data augmentation has limited ability to enhance results.
Conclusions:
EFIDL is better suited for evaluating the prediction performance of machine learning methods when data are augmented. Using unsuitable evaluation frameworks can lead to false results. Future researchers should consider the evaluation framework we proposed when working with augmented datasets. Our experiments showed that data augmentation does not significantly improve machine learning prediction performance.
Citation
Request queued. Please wait while the file is being generated. It may take some time.
Copyright
© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.