Maintenance Notice

Due to necessary scheduled maintenance, the JMIR Publications website will be unavailable from Wednesday, July 01, 2020 at 8:00 PM to 10:00 PM EST. We apologize in advance for any inconvenience this may cause you.

Who will be affected?

Accepted for/Published in: JMIR Medical Informatics

Date Submitted: Oct 3, 2019
Date Accepted: Dec 27, 2019

The final, peer-reviewed published version of this preprint can be found here:

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies

Reiner Benaim A, Almog R, Gorelik Y, Hochberg I, Nassar L, Mashiach T, Khamaisi M, Lurie Y, Azzam ZS, Khoury J, Kurnik D, Beyar R

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies

JMIR Med Inform 2020;8(2):e16492

DOI: 10.2196/16492

PMID: 32130148

PMCID: 7059086

A Validation Study for Medical Research Based on Synthetic Hospital Data

  • Anat Reiner Benaim; 
  • Ronit Almog; 
  • Yuri Gorelik; 
  • Irit Hochberg; 
  • Laila Nassar; 
  • Tanya Mashiach; 
  • Mogher Khamaisi; 
  • Yael Lurie; 
  • Zaher S. Azzam; 
  • Johad Khoury; 
  • Daniel Kurnik; 
  • Rafael Beyar

ABSTRACT

Background:

Privacy restrictions limit access to protected patient-derived health information for research purposes. Consequently, data anonymization is required to allow researchers data access for initial analysis before granting Institutional Review Board approval. A system implemented in our institution enables synthetic data generation that mimics data from real electronic medical records, wherein only fictitious patients are listed.

Objective:

This paper studies the validity of results obtained when analyzing synthetic data for medical research. A comprehensive validation process concerning meaningful clinical questions and various types of data was conducted to assess the accuracy and precision of statistical estimates derived from synthetic patient data.

Methods:

A cross-hospital project was conducted to validate results obtained from synthetic data produced for five contemporary studies on various topics. For each study, results derived from synthetic data were compared to those based on real data. In addition, repeatedly generated synthetic data sets were used to estimate the bias and stability of results obtained from synthetic data.

Results:

This study demonstrated that results derived from synthetic data were predictive of results from real data. When the number of patients was large relative to the number of variables used, highly accurate and strongly consistent results were observed between synthetic and real data. When small populations were accounted for, prediction was of moderate accuracy.

Conclusions:

The use of synthetic data provides a close estimate to real data results and is thus a powerful tool in shaping research hypotheses and accessing estimated analyses, without risking patient privacy. Synthetic data enables broad access to data, including for out-of-organization researchers, and rapid, safe, and repeatable analysis of data in hospitals or other health organizations where patient privacy is a primary value.


 Citation

Please cite as:

Reiner Benaim A, Almog R, Gorelik Y, Hochberg I, Nassar L, Mashiach T, Khamaisi M, Lurie Y, Azzam ZS, Khoury J, Kurnik D, Beyar R

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies

JMIR Med Inform 2020;8(2):e16492

DOI: 10.2196/16492

PMID: 32130148

PMCID: 7059086

Download PDF


Request queued. Please wait while the file is being generated. It may take some time.

© The authors. All rights reserved. This is a privileged document currently under peer-review/community review (or an accepted/rejected manuscript). Authors have provided JMIR Publications with an exclusive license to publish this preprint on it's website for review and ahead-of-print citation purposes only. While the final peer-reviewed paper may be licensed under a cc-by license on publication, at this stage authors and publisher expressively prohibit redistribution of this draft paper other than for review purposes.