Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Follow the data: tracking data quality and completeness in oncology real-world data.

  • TL;DR
  • Abstract
  • Literature Map
  • Similar Papers
TL;DR

This study evaluated data quality in oncology real-world data by comparing demographic and biomarker information across EHR, cancer registry, and FHIR extracts for lung cancer patients. While demographics showed high concordance, biomarker data were missing in 80-100% of FHIR extracts, revealing significant data loss and emphasizing the need for validation in RWD use.

Abstract
Translate article icon Translate Article Star icon

Electronic Health Record (EHR) data are increasingly used in cancer research, yet the fidelity of this data when exchanged between systems remains poorly quantified. This study investigated the agreement in essential biomarker data after they are passed from the EHR into the cancer registry and Fast Healthcare Interoperability Resources (FHIR) extracts. This single-institution retrospective study compared demographics and 6 biomarkers from 30 lung cancer patients seen between July 2020 and July 2022. Manual review from the EHR served as the gold standard, with concordance tested between the source EHR, Institutional Cancer Registry, and FHIR exports. Demographics showed high concordance across databases. In contrast, biomarker data present in the source EHR were missing in 80%-100% of FHIR extracts. The demographic registry variables were highly concordant. This study reports a significant loss in biomarker data availability across real-world data (RWD) sources. Results underscore critical gaps in RWD extraction or exchange methods and highlight risks of relying on RWD without validation.

Similar Papers
  • Research Article
  • Cite Count Icon 14
  • 10.1177/17407745221114298
Fitness of real-world data for clinical trial data collection: Results and lessons from a HARMONY Outcomes ancillary study.
  • Jul 24, 2022
  • Clinical Trials
  • Bradley G Hammill + 9 more

Despite the extensive use of real-world data for retrospective, observational clinical research, our understanding of how real-world data might increase the efficiency of data collection in patient-level randomized clinical trials is largely unknown. The structure of real-world data is inherently heterogeneous, with each source electronic health record and claims database different from the next. Their fitness-for-use as data sources for multisite trials in the United States has not been established. For a subset of participants in the HARMONY Outcomes Trial, we obtained electronic health record data from recruiting sites or Medicare claims data from the Centers for Medicare & Medicaid Services. For baseline characteristics and follow-up events, we assessed the level of agreement between these real-world data and data documented in the trial database. Real-world data-derived demographic information tended to agree with trial-reported demographic information, although real-world data were less accurate in identifying medical history. The ability of real-world data to identify baseline medication usage differed by real-world data source, with claims data demonstrating substantially better performance than electronic health record data. The limited number of lab results in the collected electronic health record data matched closely with values in the trial database. There were not enough follow-up events in the ancillary study population to draw meaningful conclusions about the performance of real-world data for identification of events. Based on the conduct of this ancillary study, the challenges and opportunities of using real-world data within clinical trials are discussed. Based on a subset of participants from the HARMONY Outcomes Trial, our results suggest that electronic health record or claims data, as currently available, are unlikely to be a complete substitute for trial data collection of medical history or baseline lab results, but that Medicare claims were able to identify most medications. The limited size of the study population prevents us from drawing strong conclusions based on these results, and other studies are clearly needed to confirm or refute these findings.

  • Front Matter
  • Cite Count Icon 4
  • 10.2217/cer-2021-0166
Learning from the past to advance tomorrow's real-world evidence: what demonstration projects have to teach us.
  • Sep 14, 2021
  • Journal of Comparative Effectiveness Research
  • Ashley Jaksa + 1 more

Learning from the past to advance tomorrow's real-world evidence: what demonstration projects have to teach us.

  • Research Article
  • Cite Count Icon 1
  • 10.1097/mlr.0000000000002236
Investigating the Use of the Fast Health Care Interoperability Resources (FHIR) Standard to Support Data Activities Across the PCORnet® Infrastructure: Lessons Learned From the FHIR Pilots of the Coordinating Center for PCORnet®
  • Jan 8, 2026
  • Medical Care
  • Keith Marsolo + 12 more

Background:Institutions that participate in PCORnet® transform their local electronic health record (EHR) data into the PCORnet® Common Data Model (CDM), which is then used to generate data extracts for PCORnet® Studies. PCORnet® Studies can also include institutions that do not participate in PCORnet, and for these organizations, the cost of instantiating a PCORnet® CDM can be prohibitive. Fast Health care Interoperability Resources (FHIR) provides an alternative method of obtaining EHR data.Objective:To determine whether data obtained through FHIR might be a viable study solution for those sites that do not participate in PCORnet.® This mixed-methods project had 2 objectives: (1) survey sites participating in PCORnet on the availability of FHIR (FHIR survey); (2) compare the coverage of a FHIR-based data extract using REDCap with one from the PCORnet® CDM across 3 sites (FHIR extract).Methods:(1) FHIR survey: A series of questions were asked about the use of FHIR in a production capacity. (2) FHIR extract: REDCap FHIR and PCORnet® CDM extracts were created based on study variables from 2 prior PCORnet® Studies. Data were extracted for 40 patients and concordance measures were computed between the 2 sources.Results:(1) FHIR survey: Of responding organizations, 73% (n=49) reported that FHIR was deployed in a production capacity. (2) FHIR extract: Results were highly variable. Cohen kappa ranged from 0.01 to 0.76 for certain diagnoses, 0.24 to 0.84 for laboratory results, and 0.1 to 0.87 for medications.Conclusions:Despite differences in data, certain studies may be well-suited for FHIR-based extracts.

  • Research Article
  • Cite Count Icon 2
  • 10.2196/68171
Leveraging Interoperable Electronic Health Record (EHR) Data for Distributed Analyses in Clinical Research: Technical Implementation Report of the HELP Study
  • Jul 30, 2025
  • JMIR Medical Informatics
  • Julia Palm + 3 more

BackgroundThe Medical Informatics Initiative (MII) Germany established 38 data integration centers (DIC) in university hospitals to improve health care and biomedical research through the use of electronic health record (EHR) data. To showcase the value of these DIC, the HELP (Hospital-wide Electronic Medical Record Evaluated Computerized Decision Support System to Improve Outcomes of Patients with Staphylococcal Bloodstream Infection) study was initiated as a use case. This study is a clinical trial designed to assess the impact of a computerized decision support system for managing staphylococcal bacteremia.ObjectiveIn this paper, we present the lessons learned during the use case from a technical perspective. This paper outlines the challenges encountered and solutions developed during our initial implementation of this infrastructure, providing insights applicable to other research platforms using EHR data. These insights are organized into 3 key areas: study-specific data definition and modeling, interoperable data integration and transformation, and distributed data extraction and analysis.MethodsAn interdisciplinary team of clinicians, computer scientists, and statisticians created a catalog of items to identify data elements necessary for the study’s evaluation and developed a domain-specific information model. DIC developed extract-transform-load pipelines to collect the disparate, site-specific EHR data and to transform it into a common data format. Health Level Seven International (HL7) Fast Healthcare Interoperability Resources (FHIR) and the MII’s core dataset profiles were adopted for consistent data representation across sites. Additionally, data not present in EHRs was gathered using structured electronic case report forms. Analysis scripts were then distributed to the sites to preprocess the data locally, followed by a central analysis of the preprocessed data to generate the final overall results.Implementation (Results)Our analysis revealed significant heterogeneity in data quality and implementation of interoperability standards, requiring substantial harmonization efforts. The development of analysis scripts and data extraction processes demanded multiple iterative cycles and close collaboration with local data experts. Despite these challenges, the successful implementation demonstrated the feasibility of distributed EHR analyses while highlighting the importance of thorough data quality assessment, realistic timeline planning, and multidisciplinary expertise.ConclusionsThe HELP study highlights challenges and opportunities in leveraging EHR data for clinical research, particularly in the absence of mandatory data standards and resource-intensive data harmonization efforts. Despite limitations in data availability and quality, progress in digitization and interoperability frameworks offers hope for future improvements. Lessons learned from this study can inform the development of standardized methodologies and infrastructures for sustainable EHR data integration in research.

  • Discussion
  • Cite Count Icon 71
  • 10.1016/j.jbi.2021.103871
REDCap on FHIR: Clinical Data Interoperability Services
  • Jul 21, 2021
  • Journal of Biomedical Informatics
  • A.C Cheng + 7 more

REDCap on FHIR: Clinical Data Interoperability Services

  • Research Article
  • Cite Count Icon 17
  • 10.1055/s-0042-1760436
“fhircrackr”: An R Package Unlocking Fast Healthcare Interoperability Resources for Statistical Analysis
  • Jan 1, 2023
  • Applied Clinical Informatics
  • Julia Palm + 3 more

Background The growing interest in the secondary use of electronic health record (EHR) data has increased the number of new data integration and data sharing infrastructures. The present work has been developed in the context of the German Medical Informatics Initiative, where 29 university hospitals agreed to the usage of the Health Level Seven Fast Healthcare Interoperability Resources (FHIR) standard for their newly established data integration centers. This standard is optimized to describe and exchange medical data but less suitable for standard statistical analysis which mostly requires tabular data formats.Objectives The objective of this work is to establish a tool that makes FHIR data accessible for standard statistical analysis by providing means to retrieve and transform data from a FHIR server. The tool should be implemented in a programming environment known to most data analysts and offer functions with variable degrees of flexibility and automation catering to users with different levels of FHIR expertise.Methods We propose the fhircrackr framework, which allows downloading and flattening FHIR resources for data analysis. The framework supports different download and authentication protocols and gives the user full control over the data that is extracted from the FHIR resources and transformed into tables. We implemented it using the programming language R [1] and published it under the GPL-3 open source license.Results The framework was successfully applied to both publicly available test data and real-world data from several ongoing studies. While the processing of larger real-world data sets puts a considerable burden on computation time and memory consumption, those challenges can be attenuated with a number of suitable measures like parallelization and temporary storage mechanisms.Conclusion The fhircrackr R package provides an open source solution within an environment that is familiar to most data scientists and helps overcome the practical challenges that still hamper the usage of EHR data for research.

  • Research Article
  • 10.1182/blood-2025-8138
Enhancing real-world data (RWD) quality in oncology: Implementation of the qcard initiative in a large US oncology electronic health record (EHR)-derived database for research and regulatory decision-making
  • Nov 3, 2025
  • Blood
  • Zhaohui Su + 6 more

Enhancing real-world data (RWD) quality in oncology: Implementation of the qcard initiative in a large US oncology electronic health record (EHR)-derived database for research and regulatory decision-making

  • PDF Download Icon
  • Research Article
  • 10.3390/informatics11030042
The Mappability of Clinical Real-World Data of Patients with Melanoma to Oncological Fast Healthcare Interoperability Resources (FHIR) Profiles: A Single-Center Interoperability Study
  • Jun 28, 2024
  • Informatics
  • Jessica Swoboda + 7 more

(1) Background: Tumor-specific standardized data are essential for AI-based progress in research, e.g., for predicting adverse events in patients with melanoma. Although there are oncological Fast Healthcare Interoperability Resources (FHIR) profiles, it is unclear how well these can represent malignant melanoma. (2) Methods: We created a methodology pipeline to assess to what extent an oncological FHIR profile, in combination with a standard FHIR specification, can represent a real-world data set. We extracted Electronic Health Record (EHR) data from a data platform, and identified and validated relevant features. We created a melanoma data model and mapped its features to the oncological HL7 FHIR Basisprofil Onkologie [Basic Profile Oncology] and the standard FHIR specification R4. (3) Results: We identified 216 features. Mapping showed that 45 out of 216 (20.83%) features could be mapped completely or with adjustments using the Basisprofil Onkologie [Basic Profile Oncology], and 129 (60.85%) features could be mapped using the standard FHIR specification. A total of 39 (18.06%) new, non-mappable features could be identified. (4) Conclusions: Our tumor-specific real-world melanoma data could be partially mapped using a combination of an oncological FHIR profile and a standard FHIR specification. However, important data features were lost or had to be mapped with self-defined extensions, resulting in limited interoperability.

  • Research Article
  • Cite Count Icon 23
  • 10.1016/j.cmpb.2021.106232
Development of an application concerning fast healthcare interoperability resources based on standardized structured medical information exchange version 2 data
  • Jun 8, 2021
  • Computer Methods and Programs in Biomedicine
  • Dingding Xiao + 3 more

Development of an application concerning fast healthcare interoperability resources based on standardized structured medical information exchange version 2 data

  • Conference Article
  • Cite Count Icon 10
  • 10.1145/3301879.3301881
Implementation of SMART on FHIR in Developing Countries Through SFPBRF
  • Nov 12, 2018
  • Abrar Ahmad + 2 more

Fast Healthcare Interoperability Resources (FHIR) is an International health standard for health data developed by Health Level Seven -- (HL7) an international organization for the development of health data standards. FHIR enables data interoperability as it is based on lightweight open source RESTful services. Many developed countries which already running electronic health record systems are now focused on the adoption of FHIR to unlock its potential benefits by integrating other technologies like Substitutable Medical Applications, Reusable Technology (SMART) platform. In most of the developing countries electronic health record system does not exist because they lag in resources to invest in electronic health record systems and they are still operating on paper-based health record system, so they are unable to adopt FHIR and other related technologies like SMART. Due to which, interoperability is not enabled on this paper-based data. These countries remain at a distance from the benefits of FHIR and SMART platform to provide better patient care to their patients, quick and efficient clinical decision making and better diagnostics. This paper presents the implementation of SMART on FHIR in healthcare organizations of developing countries through a proposed framework SFPBRF which not only maps paper-based health record system's data to HL7's FHIR standard but also integrate the complete SMART on FHIR platform to run SMART apps on this FHIR conformed data. This paper presents successful translation (done by translation engine) of paper-based data to HL7 FHIR standard. It also shows the running of open source and internally developed SMART apps on this FHIR conformed data. Thus, by the successful mapping and implementation of SMART on FHIR through proposed framework SFPBRF we can conclude that FHIR can be adopted in healthcare organizations of developing countries and SMART on FHIR can help a lot in achieving better patient care, quick and efficient decision making and better diagnostics in developing countries.

  • Research Article
  • 10.1093/jamia/ocag095
A comparison of Fast Healthcare Interoperability Resources and Observational Medical Mutcomes Partnership electronic health record data within the All of Us Research Program.
  • Jun 23, 2026
  • Journal of the American Medical Informatics Association : JAMIA
  • Jason Patterson + 7 more

This study compares the contents of two data standards; the Observational Medical Outcomes Partnership (OMOP) and Fast Healthcare Interoperability Resources (FHIR), highlighting their strength and weaknesses and serve as an initial step toward understanding how each standard supports secondary data analysis. Participant electronic health record data in both OMOP and FHIR formats from the All of Us Research Program (AoURP) were compared, including codeable event volume, healthcare encounters, and person timelines. A phenotype-based assessment was also conducted using Type-II Diabetes Mellitus (T2DM). Among 29512 participants identified with overlapping FHIR and OMOP data, Median codeable event counts were comparable between FHIR and OMOP within the Measurement (OMOP = 846; FHIR = 832), Drug (OMOP = 92; FHIR = 90), and Observation (OMOP = 65; FHIR = 100) domains, but were higher in OMOP within the Condition (OMOP = 258; FHIR = 11) and Procedure (OMOP = 72; FHIR = 4) domains. Within the T2DM cohort, OMOP contained more data, except for medications. Very few participants had encounters in FHIR (1.5%) relative to OMOP (97.0%). On average, only 15.9% of visit dates overlapped in both standards, with most visit dates occurring only in OMOP (65.3%) or only FHIR (18.4%). Data in OMOP showed a higher volume of observation and procedure codeable events, reported encounters, and T2DM symptoms, complications, and comorbidities. FHIR data was able to capture data across multiple providers and health systems. Based on AoURP data, both standards were shown to support healthcare data capture, although OMOP shows greater utility for research purposes as its extract, transform, and load process enables more flexible data capture relative to extracting data from FHIR payloads. FHIR, however, captures patient-level data from beyond the health system and can thus be used to supplement OMOP.

  • Research Article
  • Cite Count Icon 57
  • 10.1093/jamiaopen/ooz056
Developing a scalable FHIR-based clinical data normalization pipeline for standardizing and integrating unstructured and structured electronic health record data.
  • Oct 18, 2019
  • JAMIA Open
  • Na Hong + 6 more

ObjectiveTo design, develop, and evaluate a scalable clinical data normalization pipeline for standardizing unstructured electronic health record (EHR) data leveraging the HL7 Fast Healthcare Interoperability Resources (FHIR) specification.MethodsWe established an FHIR-based clinical data normalization pipeline known as NLP2FHIR that mainly comprises: (1) a module for a core natural language processing (NLP) engine with an FHIR-based type system; (2) a module for integrating structured data; and (3) a module for content normalization. We evaluated the FHIR modeling capability focusing on core clinical resources such as Condition, Procedure, MedicationStatement (including Medication), and FamilyMemberHistory using Mayo Clinic’s unstructured EHR data. We constructed a gold standard reusing annotation corpora from previous NLP projects.ResultsA total of 30 mapping rules, 62 normalization rules, and 11 NLP-specific FHIR extensions were created and implemented in the NLP2FHIR pipeline. The elements that need to integrate structured data from each clinical resource were identified. The performance of unstructured data modeling achieved F scores ranging from 0.69 to 0.99 for various FHIR element representations (0.69–0.99 for Condition; 0.75–0.84 for Procedure; 0.71–0.99 for MedicationStatement; and 0.75–0.95 for FamilyMemberHistory).ConclusionWe demonstrated that the NLP2FHIR pipeline is feasible for modeling unstructured EHR data and integrating structured elements into the model. The outcomes of this work provide standards-based tools of clinical data normalization that is indispensable for enabling portable EHR-driven phenotyping and large-scale data analytics, as well as useful insights for future developments of the FHIR specifications with regard to handling unstructured clinical data.

  • Research Article
  • Cite Count Icon 61
  • 10.1055/s-0038-1676466
A Pharmacogenomics Clinical Decision Support Service Based on FHIR and CDS Hooks.
  • Dec 1, 2018
  • Methods of Information in Medicine
  • R.H Dolin + 2 more

Pharmacogenomics (PGx) is often considered a low-hanging fruit for genomics-electronic health record (EHR) integrations, and many have expressed the notion that drug-gene interaction checking might one day become as much a commodity in EHRs as drug-drug and drug-allergy checking. In addition, the U.S. Office of the National Coordinator has recognized the trend toward storing complete sequencing data outside the EHR in a Genomic Archiving and Communication System (GACS) and has emphasized the need for "pilots that test Fast Healthcare Interoperability Resources (FHIR) Genomics for GACS integration with EHRs." We sought to develop a PGx clinical decision support (CDS) service, leveraging the emerging FHIR and CDS Hooks standards, and based on an assumption that pharmacogene sequencing data would be stored alongside the EHR in a GACS. We developed a PGx CDS service as a functional prototype. The service is triggered by a medication order in the EHR. When evoked, the service looks for relevant genetic data in a GACS and returns corresponding recommendations back to the ordering clinician. Where the patient has no genetic data on file, the service can recommend pretreatment genetic testing where applicable. Overall, we were able to meet our objectives and deploy a functional prototype, interfaced with a commercial EHR. We identified several areas where FHIR or CDS Hooks lacked necessary semantics or have implementation ambiguity. Primary FHIR challenges included multiple ways to say the same thing, which exacerbated the complexity of variant to allele conversion and lack of representation of deoxyribonucleic acid region(s) studied. Primary CDS Hooks challenges included the complexity of executing an authenticated query against one system (GACS) upon being triggered by a different system (the EHR), and limitations in the types of actionable recommendations that can be returned to the EHR. In conclusion, we have found that PGx CDS based on FHIR and CDS Hooks appears to represent a promising means of genomics-EHR integration. More real-world testing along with a set of use-case driven GACS interface requirements will push us closer to the U.S. National Human Genome Research Institute vision of a plug-in PGx app.

  • Research Article
  • Cite Count Icon 3
  • 10.1200/cci.24.00100
Automated Electronic Health Record Data Extraction and Curation Using ExtractEHR
  • Nov 1, 2024
  • JCO Clinical Cancer Informatics
  • Tamara P Miller + 10 more

PURPOSEAlthough the potential transformative effect of electronic health record (EHR) data on clinical research in adult patient populations has been very extensively discussed, the effect on pediatric oncology research has been limited. Multiple factors contribute to this more limited effect, including the paucity of pediatric cancer cases in commercial EHR-derived cancer data sets and phenotypic case identification challenges in pediatric federated EHR data.METHODSThe ExtractEHR software package was initially developed as a tool to improve clinical trial adverse event reporting but has expanded its use cases to include the development of multisite EHR data sets and the support of cancer cohorts. ExtractEHR enables customized, automated data extraction from the EHR that, when implemented across multiple hospitals, can create pediatric cancer EHR data sets to address a very wide range of research questions in pediatric oncology. After ExtractEHR data acquisition, EHR data can be cleaned and graded using CleanEHR and GradeEHR, companion software packages.RESULTSExtractEHR has been installed at four leading pediatric institutions: Children's Healthcare of Atlanta, Children's Hospital of Philadelphia, Texas Children's Hospital, and Seattle Children's Hospital.CONCLUSIONExtractEHR has supported multiple use cases, including five clinical epidemiology studies, multicenter clinical trials, and cancer cohort assembly. Work is ongoing to develop Fast Health care Interoperability Resources ExtractEHR and implement other sustainability and scalability enhancements.

  • Research Article
  • Cite Count Icon 95
  • 10.3233/shti190805
The Use of FHIR in Digital Health - A Review of the Scientific Literature.
  • Jan 1, 2019
  • Studies in health technology and informatics
  • Lehne Moritz + 3 more

Fast Healthcare Interoperability Resources (FHIR), an international standard for exchanging digital health data, is increasingly used in health information technology. FHIR promises to facilitate the use of electronic health records (EHRs), enable mobile technologies and make health data accessible to large-scale analytics. Until now, there is no comprehensive review of scientific articles about FHIR and its use in digital health. Here, we aim to address this gap and provide an overview of the main topics associated with FHIR in the scientific literature. For this, we screened all articles about FHIR on Web of Science and PubMed and identified the main topics discussed in these articles. We also explored the temporal trend and geography of publications and performed some basic text mining on article abstracts. We found that the topics most commonly discussed in the articles were related to data models, mobile and web applications as well as medical devices. Since its introduction, the number of publications about FHIR have steadily increased until 2017, indicating an increasing popularity of FHIR in healthcare (in 2018, publication numbers remained stable). In sum, our study provides an overview of the scientific literature about FHIR and its current use in digital health.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant