Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies.

Anat Reiner Benaim,Yael Lurie,Rafael Beyar,Tanya Mashiach,Ronit Almog,Yuri Gorelik,Laila Nassar,Mogher Khamaisi,Daniel Kurnik,Zaher S Azzam,Irit Hochberg,Johad Khoury

doi:10.2196/16492

Anat Reiner Benaim, Yael Lurie + Show 10 more

Open Access

PDF Available

https://doi.org/10.2196/16492

Copy DOI

Export

Save

Cite

Abstract
Highlights/Summary
Full-Text PDF
Similar Papers

Abstract

Listen

BackgroundPrivacy restrictions limit access to protected patient-derived health information for research purposes. Consequently, data anonymization is required to allow researchers data access for initial analysis before granting institutional review board approval. A system installed and activated at our institution enables synthetic data generation that mimics data from real electronic medical records, wherein only fictitious patients are listed.ObjectiveThis paper aimed to validate the results obtained when analyzing synthetic structured data for medical research. A comprehensive validation process concerning meaningful clinical questions and various types of data was conducted to assess the accuracy and precision of statistical estimates derived from synthetic patient data.MethodsA cross-hospital project was conducted to validate results obtained from synthetic data produced for five contemporary studies on various topics. For each study, results derived from synthetic data were compared with those based on real data. In addition, repeatedly generated synthetic datasets were used to estimate the bias and stability of results obtained from synthetic data.ResultsThis study demonstrated that results derived from synthetic data were predictive of results from real data. When the number of patients was large relative to the number of variables used, highly accurate and strongly consistent results were observed between synthetic and real data. For studies based on smaller populations that accounted for confounders and modifiers by multivariate models, predictions were of moderate accuracy, yet clear trends were correctly observed.ConclusionsThe use of synthetic structured data provides a close estimate to real data results and is thus a powerful tool in shaping research hypotheses and accessing estimated analyses, without risking patient privacy. Synthetic data enable broad access to data (eg, for out-of-organization researchers), and rapid, safe, and repeatable analysis of data in hospitals or other health organizations where patient privacy is a primary value.

Highlights

BackgroundAccess to large databases of electronic medical records (EMRs) for research purposes is limited by privacy restriction, security laws and regulations, and organizational guidelines imposed because of the assumed value of the data
This study demonstrated that results derived from synthetic data were predictive of results from real data
Between 2007 and 2017, we identified 12,188 patients discharged on oral anticoagulant Observational Medical Dataset Simulator (OSIM) (OAC), some of whom received a single antiplatelet, either aspirin (n=3953) or P2Y ADP receptor blockers antiplatelet therapy (n=882), or a double antiplatelet therapy (DAT; n=417)

Summary

Introduction

Access to large databases of electronic medical records (EMRs) for research purposes is limited by privacy restriction, security laws and regulations, and organizational guidelines imposed because of the assumed value of the data. Requires approval of the local institutional review board (IRB), but this regulatory process is often time consuming, thereby delaying research and imposing difficulties on data sharing and collaborations. Data anonymization, namely, making reidentification of patients impossible, is required to balance the risk of privacy intrusions with research accessibility. Data anonymization is required to allow researchers data access for initial analysis before granting institutional review board approval. A system installed and activated at our institution enables synthetic data generation that mimics data from real electronic medical records, wherein only fictitious patients are listed

Methods

Results

Discussion

Conclusion

Full Text

Published Version (Free)

View/Download pdf

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: JMIR Medical Informatics	Publication Date: Feb 20, 2020
Citations: 90	License type: cc-by

R Discovery Prime

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies.

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: JMIR Medical Informatics

Lead the way for us

Similar Papers

Weighted Itemsets Error (WIE) Approach for Evaluating Generated Synthetic Patient Data
Mojtaba Zare ... Janusz Wojtusiak
-
Mojtaba Zare, et. al.Mojtaba Zare ... Janusz Wojtusiak
01 Dec 2018
01 Dec 2018

Systematic Evaluation of Synthetic Panel Data Quality with an Application to Chronic Lymphocytic Leukemia
Dimitris Karletsos ... Andy Wilson
Blood | VOL. 140
Dimitris Karletsos, et. al.Dimitris Karletsos ... Andy Wilson
15 Nov 2022
Blood | VOL. 140

The potential synergies between synthetic data and in silico trials in relation to generating representative virtual population cohorts
Puja Myles ... Johan Ordish
Progress in Biomedical Engineering | VOL. 5
Puja Myles, et. al.Puja Myles ... Johan Ordish
01 Jan 2023
Progress in Biomedical Engineering | VOL. 5

Synthetic Data Generation By Artificial Intelligence to Accelerate Translational Research and Precision Medicine in Hematological Malignancies
Saverio D'Amico ...
Blood | VOL. 140
Saverio D'Amico, et. al.Saverio D'Amico ...
15 Nov 2022
Blood | VOL. 140

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

Analyzing Medical Research Results Based on Synthetic Data and Their Relation to Real Data Results: Systematic Comparison From Five Observational Studies.

Abstract

Highlights

Summary

Published Version (Free)

Talk to us

Similar Papers

More From: JMIR Medical Informatics