Evaluating current automatic de-identification methods with Veteran’s health administration clinical documents

Oscar Ferrández,Stéphane M Meystre,Shuying Shen,Brett R South,F Jeffrey Friedlin,Matthew H Samore

doi:10.1186/1471-2288-12-109

Abstract

BackgroundThe increased use and adoption of Electronic Health Records (EHR) causes a tremendous growth in digital information useful for clinicians, researchers and many other operational purposes. However, this information is rich in Protected Health Information (PHI), which severely restricts its access and possible uses. A number of investigators have developed methods for automatically de-identifying EHR documents by removing PHI, as specified in the Health Insurance Portability and Accountability Act “Safe Harbor” method.This study focuses on the evaluation of existing automated text de-identification methods and tools, as applied to Veterans Health Administration (VHA) clinical documents, to assess which methods perform better with each category of PHI found in our clinical notes; and when new methods are needed to improve performance.MethodsWe installed and evaluated five text de-identification systems “out-of-the-box” using a corpus of VHA clinical documents. The systems based on machine learning methods were trained with the 2006 i2b2 de-identification corpora and evaluated with our VHA corpus, and also evaluated with a ten-fold cross-validation experiment using our VHA corpus. We counted exact, partial, and fully contained matches with reference annotations, considering each PHI type separately, or only one unique ‘PHI’ category. Performance of the systems was assessed using recall (equivalent to sensitivity) and precision (equivalent to positive predictive value) metrics, as well as the F2-measure.ResultsOverall, systems based on rules and pattern matching achieved better recall, and precision was always better with systems based on machine learning approaches. The highest “out-of-the-box” F2-measure was 67% for partial matches; the best precision and recall were 95% and 78%, respectively. Finally, the ten-fold cross validation experiment allowed for an increase of the F2-measure to 79% with partial matches.ConclusionsThe “out-of-the-box” evaluation of text de-identification systems provided us with compelling insight about the best methods for de-identification of VHA clinical documents. The errors analysis demonstrated an important need for customization to PHI formats specific to VHA documents. This study informed the planning and development of a “best-of-breed” automatic de-identification application for VHA clinical text.

Highlights

With the increased use and adoption of Electronic Health Records (EHR) systems, we have witnessed a tremendous growth in digital information useful for clinicians, researchers and many other operational purposes
Principal methods used for automatic de-identification many systems combine different approaches to de-identify specific Protected Health Information (PHI) types, we present here a broad classification depending on the main technique used to obscure PHI
We have presented a study about the suitability of current text de-identification methods and tools for deidentifying Veterans Health Administration (VHA) clinical documents

Summary

Introduction

The increased use and adoption of Electronic Health Records (EHR) causes a tremendous growth in digital information useful for clinicians, researchers and many other operational purposes. With the increased use and adoption of Electronic Health Records (EHR) systems, we have witnessed a tremendous growth in digital information useful for clinicians, researchers and many other operational purposes This vastness of data is rich in Protected Health Information (PHI), which severely restricts its access and possible uses. In the United States, the confidentiality of patient data is protected by the Health Insurance Portability and Accountability Act (HIPAA; codified as 45 CFR }160 and 164) and the Common Rule [1] These laws typically require the informed consent of the patient and approval of the Internal Review Board (IRB) to use data for research purposes, but these requirements are sometimes extremely difficult or even impossible to fulfill (e.g., retrospective studies of large patient populations who moved, changed healthcare system, or died). For clinical data to be considered de-identified, the HIPAA “Safe Harbor” technique requires 18 PHI identifiers to be removed; further details regarding these 18 PHI identifiers can be found in [2,3]

Methods

Results

Conclusion

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: BMC Medical Research Methodology	Publication Date: Jul 27, 2012
Citations: 57	License type: CC BY 2.0

R Discovery Prime

R Discovery Prime

Evaluating current automatic de-identification methods with Veteran’s health administration clinical documents

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Medical Research Methodology

Lead the way for us

Similar Papers

Local public health department adoption and use of electronic health records.
J Mac Mccullough ... Douglas S Bell
Journal of Public Health Management and Practice | VOL. 21
J Mac Mccullough, et. al.J Mac Mccullough ... Douglas S Bell
01 Jan 2015
Journal of Public Health Management and Practice | VOL. 21

Automatic de-identification of textual documents in the electronic health record: a review of recent research
Stephane M Meystre ... Shuying Shen
BMC Medical Research Methodology | VOL. 10
Stephane M Meystre, et. al.Stephane M Meystre ... Shuying Shen
02 Aug 2010
BMC Medical Research Methodology | VOL. 10

Electronic Health Records: Promises and Realities: Part III: Information Privacy and Accuracy: Zero and GIGO Won't Do
William B Millard
Annals of Emergency Medicine | VOL. 56
William B MillardWilliam B Millard
22 Sep 2010
Annals of Emergency Medicine | VOL. 56

Evolution of "the guideline advantage": lessons learned from the front lines of outpatient performance measurement.
Vincent Bufalino ... Jay H Shubrook
Circulation. Cardiovascular quality and outcomes | VOL. 7
Vincent Bufalino, et. al.Vincent Bufalino ... Jay H Shubrook
30 Apr 2014
Circulation. Cardiovascular quality and outcomes | VOL. 7

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Evaluating current automatic de-identification methods with Veteran’s health administration clinical documents

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: BMC Medical Research Methodology