Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Unfolding Data Quality Dimensions in Practice: A Survey

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Data quality can be assessed across multiple high-level concepts called dimensions , such as accuracy, completeness, consistency, and timeliness. While extensive research and several attempts for standardization (e.g., ISO/IEC 25012) exist for data quality dimensions, their practical application often remains unclear. In parallel to research endeavors, a large number of tools have been developed that implement functionalities for the detection and mitigation of specific data quality issues, such as missing values or outliers. With this paper, we aim to bridge this gap between data quality theory and practice by systematically connecting low-level functionalities offered by data quality tools with high-level dimensions, revealing their many-to-many relationships. Through an examination of seven open-source data quality tools, we provide a comprehensive mapping between their functionalities and the data quality dimensions, demonstrating how individual functionalities and their variants partially contribute to the assessment of single dimensions. This systematic survey provides both practitioners and researchers with a unified view on the fragmented landscape of data quality checks, offering actionable insights for quality assessment across multiple dimensions.

Similar Papers
  • Research Article
  • Cite Count Icon 120
  • 10.1093/jamia/ocad120
Electronic health record data quality assessment and tools: a systematic review.
  • Jun 30, 2023
  • Journal of the American Medical Informatics Association : JAMIA
  • Abigail E Lewis + 6 more

We extended a 2013 literature review on electronic health record (EHR) data quality assessment approaches and tools to determine recent improvements or changes in EHR data quality assessment methodologies. We completed a systematic review of PubMed articles from 2013 to April 2023 that discussed the quality assessment of EHR data. We screened and reviewed papers for the dimensions and methods defined in the original 2013 manuscript. We categorized papers as data quality outcomes of interest, tools, or opinion pieces. We abstracted and defined additional themes and methods though an iterative review process. We included 103 papers in the review, of which 73 were data quality outcomes of interest papers, 22 were tools, and 8 were opinion pieces. The most common dimension of data quality assessed was completeness, followed by correctness, concordance, plausibility, and currency. We abstracted conformance and bias as 2 additional dimensions of data quality and structural agreement as an additional methodology. There has been an increase in EHR data quality assessment publications since the original 2013 review. Consistent dimensions of EHR data quality continue to be assessed across applications. Despite consistent patterns of assessment, there still does not exist a standard approach for assessing EHR data quality. Guidelines are needed for EHR data quality assessment to improve the efficiency, transparency, comparability, and interoperability of data quality assessment. These guidelines must be both scalable and flexible. Automation could be helpful in generalizing this process.

  • Research Article
  • Cite Count Icon 90
  • 10.1093/jamia/ocaa245
Assessing the practice of data quality evaluation in a national clinical data research network through a systematic scoping review in the era of real-world data.
  • Nov 9, 2020
  • Journal of the American Medical Informatics Association : JAMIA
  • Jiang Bian + 11 more

ObjectiveTo synthesize data quality (DQ) dimensions and assessment methods of real-world data, especially electronic health records, through a systematic scoping review and to assess the practice of DQ assessment in the national Patient-centered Clinical Research Network (PCORnet).Materials and MethodsWe started with 3 widely cited DQ literature—2 reviews from Chan et al (2010) and Weiskopf et al (2013a) and 1 DQ framework from Kahn et al (2016)—and expanded our review systematically to cover relevant articles published up to February 2020. We extracted DQ dimensions and assessment methods from these studies, mapped their relationships, and organized a synthesized summarization of existing DQ dimensions and assessment methods. We reviewed the data checks employed by the PCORnet and mapped them to the synthesized DQ dimensions and methods.ResultsWe analyzed a total of 3 reviews, 20 DQ frameworks, and 226 DQ studies and extracted 14 DQ dimensions and 10 assessment methods. We found that completeness, concordance, and correctness/accuracy were commonly assessed. Element presence, validity check, and conformance were commonly used DQ assessment methods and were the main focuses of the PCORnet data checks.DiscussionDefinitions of DQ dimensions and methods were not consistent in the literature, and the DQ assessment practice was not evenly distributed (eg, usability and ease-of-use were rarely discussed). Challenges in DQ assessments, given the complex and heterogeneous nature of real-world data, exist.ConclusionThe practice of DQ assessment is still limited in scope. Future work is warranted to generate understandable, executable, and reusable DQ measures.

  • PDF Download Icon
  • Supplementary Content
  • Cite Count Icon 92
  • 10.2196/42615
Digital Health Data Quality Issues: Systematic Review
  • Mar 31, 2023
  • Journal of Medical Internet Research
  • Rehan Syed + 13 more

BackgroundThe promise of digital health is principally dependent on the ability to electronically capture data that can be analyzed to improve decision-making. However, the ability to effectively harness data has proven elusive, largely because of the quality of the data captured. Despite the importance of data quality (DQ), an agreed-upon DQ taxonomy evades literature. When consolidated frameworks are developed, the dimensions are often fragmented, without consideration of the interrelationships among the dimensions or their resultant impact.ObjectiveThe aim of this study was to develop a consolidated digital health DQ dimension and outcome (DQ-DO) framework to provide insights into 3 research questions: What are the dimensions of digital health DQ? How are the dimensions of digital health DQ related? and What are the impacts of digital health DQ?MethodsFollowing the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, a developmental systematic literature review was conducted of peer-reviewed literature focusing on digital health DQ in predominately hospital settings. A total of 227 relevant articles were retrieved and inductively analyzed to identify digital health DQ dimensions and outcomes. The inductive analysis was performed through open coding, constant comparison, and card sorting with subject matter experts to identify digital health DQ dimensions and digital health DQ outcomes. Subsequently, a computer-assisted analysis was performed and verified by DQ experts to identify the interrelationships among the DQ dimensions and relationships between DQ dimensions and outcomes. The analysis resulted in the development of the DQ-DO framework.ResultsThe digital health DQ-DO framework consists of 6 dimensions of DQ, namely accessibility, accuracy, completeness, consistency, contextual validity, and currency; interrelationships among the dimensions of digital health DQ, with consistency being the most influential dimension impacting all other digital health DQ dimensions; 5 digital health DQ outcomes, namely clinical, clinician, research-related, business process, and organizational outcomes; and relationships between the digital health DQ dimensions and DQ outcomes, with the consistency and accessibility dimensions impacting all DQ outcomes.ConclusionsThe DQ-DO framework developed in this study demonstrates the complexity of digital health DQ and the necessity for reducing digital health DQ issues. The framework further provides health care executives with holistic insights into DQ issues and resultant outcomes, which can help them prioritize which DQ-related problems to tackle first.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.1177/08944393241245395
Assessing Data Quality in the Age of Digital Social Research: A Systematic Review
  • Apr 27, 2024
  • Social Science Computer Review
  • Jessica Daikeler + 9 more

While survey data has long been the focus of quantitative social science analyses, observational and content data, although long-established, are gaining renewed attention; especially when this type of data is obtained by and for observing digital content and behavior. Today, digital technologies allow social scientists to track “everyday behavior” and to extract opinions from public discussions on online platforms. These new types of digital traces of human behavior, together with computational methods for analyzing them, have opened new avenues for analyzing, understanding, and addressing social science research questions. However, even the most innovative and extensive amounts of data are hollow if they are not of high quality. But what does data quality mean for modern social science data? To investigate this rather abstract question the present study focuses on four objectives. First, we provide researchers with a decision tree to identify appropriate data quality frameworks for a given use case. Second, we determine which data types and quality dimensions are already addressed in the existing frameworks. Third, we identify gaps with respect to different data types and data quality dimensions within the existing frameworks which need to be filled. And fourth, we provide a detailed literature overview for the intrinsic and extrinsic perspectives on data quality. By conducting a systematic literature review based on text mining methods, we identified and reviewed 58 data quality frameworks. In our decision tree, the three categories, namely, data type, the perspective it takes, and its level of granularity, help researchers to find appropriate data quality frameworks. We, furthermore, discovered gaps in the available frameworks with respect to visual and especially linked data and point out in our review that even famous frameworks might miss important aspects. The article ends with a critical discussion of the current state of the literature and potential future research avenues.

  • Research Article
  • Cite Count Icon 7
  • 10.5210/ojphi.v10i2.9317
Evaluation of Data Exchange Process for Interoperability and Impact on Electronic Laboratory Reporting Quality to a State Public Health Agency.
  • Sep 21, 2018
  • Online Journal of Public Health Informatics
  • Sripriya Rajamani + 3 more

BackgroundPast and present national initiatives advocate for electronic exchange of health data and emphasize interoperability. The critical role of public health in the context of disease surveillance was recognized with recommendations for electronic laboratory reporting (ELR). Many public health agencies have seen a trend towards centralization of information technology services which adds another layer of complexity to interoperability efforts.ObjectivesThe study objective was to understand the process of data exchange and its impact on the quality of data being transmitted in the context of electronic laboratory reporting to public health. This was conducted in context of Minnesota Electronic Disease Surveillance System (MEDSS), the public health information system for supporting infectious disease surveillance in Minnesota. Data Quality (DQ) dimensions by Strong et al., was chosen as the guiding framework for evaluation.MethodsThe process of assessing data exchange for electronic lab reporting and its impact was a mixed methods approach with qualitative data obtained through expert discussions and quantitative data obtained from queries of the MEDSS system. Interviews were conducted in an open-ended format from November 2017 through February 2018. Based on these discussions, two high level categories of data exchange process which could impact data quality were identified: onboarding for electronic lab reporting and internal data exchange routing. This in turn comprised of ten critical steps and its impact on quality of data was identified through expert input. This was followed by analysis of data in MEDSS by various criteria identified by the informatics team.ResultsAll DQ metrics (Intrinsic DQ, Contextual DQ, Representational DQ, and Accessibility DQ) were impacted in the data exchange process with varying influence on DQ dimensions. Some errors such as improper mapping in electronic health records (EHRs) and laboratory information systems had a cascading effect and can pass through technical filters and go undetected till use of data by epidemiologists. Some DQ dimensions such as accuracy, relevancy, value-added data and interpretability are more dependent on users at either end of the data exchange spectrum, the relevant clinical groups and the public health program professionals. The study revealed that data quality is dynamic and on-going oversight is a combined effort by MEDSS Informatics team and review by technical and public health program professionals.ConclusionWith increasing electronic reporting to public health, there is a need to understand the current processes for electronic exchange and their impact on quality of data. This study focused on electronic laboratory reporting to public health and analyzed both onboarding and internal data exchange processes. Insights gathered from this research can be applied to other public health reporting currently (e.g. immunizations) and will be valuable in planning for electronic case reporting in near future.

  • Research Article
  • Cite Count Icon 32
  • 10.1055/s-0043-1761500
Data Quality in Health Care: Main Concepts and Assessment Methodologies.
  • Jan 30, 2023
  • Methods of Information in Medicine
  • Mehrnaz Mashoufi + 3 more

In the health care environment, a huge volume of data is produced on a daily basis. However, the processes of collecting, storing, sharing, analyzing, and reporting health data usually face with numerous challenges that lead to producing incomplete, inaccurate, and untimely data. As a result, data quality issues have received more attention than before. The purpose of this article is to provide an insight into the data quality definitions, dimensions, and assessment methodologies. In this article, a scoping literature review approach was used to describe and summarize the main concepts related to data quality and data quality assessment methodologies. Search terms were selected to find the relevant articles published between January 1, 2012 and September 31, 2022. The retrieved articles were then reviewed and the results were reported narratively. In total, 23 papers were included in the study. According to the results, data quality dimensions were various and different methodologies were used to assess them. Most studies used quantitative methods to measure data quality dimensions either in paper-based or computer-based medical records. Only two studies investigated respondents' opinions about data quality. In health care, high-quality data not only are important for patient care, but also are vital for improving quality of health care services and better decision making. Therefore, using technical and nontechnical solutions as well as constant assessment and supervision is suggested to improve data quality.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.3390/buildings13040944
Determinants of Data Quality Dimensions for Assessing Highway Infrastructure Data Using Semiotic Framework
  • Apr 2, 2023
  • Buildings
  • Chenchu Murali Krishna + 2 more

The rapid accumulation of highway infrastructure data and their widespread reuse in decision-making poses data quality issues. To address the data quality issue, it is necessary to comprehend data quality, followed by approaches for enhancing data quality and decision-making based on data quality information. This research aimed to identify the critical data quality dimensions that affect the decision-making process of highway projects. Firstly, a state-of-the-art review of data quality frameworks applied in various fields was conducted to identify suitable frameworks for highway infrastructure data. Data quality dimensions of the semiotic framework were identified from the literature, and an interview was conducted with the highway infrastructure stakeholders to finalise the data quality dimension. Then, a questionnaire survey identified the critical data quality dimensions for decision-making. Along with the critical dimensions, their level of importance was also identified at each highway infrastructure project’s decision-making levels. The semiotic data quality framework provided a theoretical foundation for developing data quality dimensions to assess subjective data quality. Further research is required to find effective ways to assess current data quality satisfaction at the decision-making levels.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 26
  • 10.2196/31618
Identifying Data Quality Dimensions for Person-Generated Wearable Device Data: Multi-Method Study.
  • Dec 23, 2021
  • JMIR mHealth and uHealth
  • Sylvia Cho + 3 more

BackgroundThere is a growing interest in using person-generated wearable device data for biomedical research, but there are also concerns regarding the quality of data such as missing or incorrect data. This emphasizes the importance of assessing data quality before conducting research. In order to perform data quality assessments, it is essential to define what data quality means for person-generated wearable device data by identifying the data quality dimensions.ObjectiveThis study aims to identify data quality dimensions for person-generated wearable device data for research purposes.MethodsThis study was conducted in 3 phases: literature review, survey, and focus group discussion. The literature review was conducted following the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guideline to identify factors affecting data quality and its associated data quality challenges. In addition, we conducted a survey to confirm and complement results from the literature review and to understand researchers’ perceptions on data quality dimensions that were previously identified as dimensions for the secondary use of electronic health record (EHR) data. We sent the survey to researchers with experience in analyzing wearable device data. Focus group discussion sessions were conducted with domain experts to derive data quality dimensions for person-generated wearable device data. On the basis of the results from the literature review and survey, a facilitator proposed potential data quality dimensions relevant to person-generated wearable device data, and the domain experts accepted or rejected the suggested dimensions.ResultsIn total, 19 studies were included in the literature review, and 3 major themes emerged: device- and technical-related, user-related, and data governance–related factors. The associated data quality problems were incomplete data, incorrect data, and heterogeneous data. A total of 20 respondents answered the survey. The major data quality challenges faced by researchers were completeness, accuracy, and plausibility. The importance ratings on data quality dimensions in an existing framework showed that the dimensions for secondary use of EHR data are applicable to person-generated wearable device data. There were 3 focus group sessions with domain experts in data quality and wearable device research. The experts concluded that intrinsic data quality features, such as conformance, completeness, and plausibility, and contextual and fitness-for-use data quality features, such as completeness (breadth and density) and temporal data granularity, are important data quality dimensions for assessing person-generated wearable device data for research purposes.ConclusionsIn this study, intrinsic and contextual and fitness-for-use data quality dimensions for person-generated wearable device data were identified. The dimensions were adapted from data quality terminologies and frameworks for the secondary use of EHR data with a few modifications. Further research on how data quality can be assessed with respect to each dimension is needed.

  • Research Article
  • Cite Count Icon 13
  • 10.3390/bdcc9040093
A Comparison of Data Quality Frameworks: A Review
  • Apr 9, 2025
  • Big Data and Cognitive Computing
  • Russell Miller + 3 more

This study reviews various data quality frameworks that have some form of regulatory backing. The aim is to identify how these frameworks define, measure, and apply data quality dimensions. This review identified generalisable frameworks, such as TDQM, ISO 8000, and ISO 25012, and specialised frameworks, such as IMF’s DQAF, BCBS 239, WHO’s DQA, and ALCOA+. A standardised data quality model was employed to map the dimensions of the data from each framework to a common vocabulary. This mapping enabled a gap analysis that highlights the presence or absence of specific data quality dimensions across the examined frameworks. The analysis revealed that core data quality dimensions such as “accuracy”, “completeness”, “consistency”, and “timeliness” are equally and well represented across all frameworks. In contrast, dimensions such as “semantics” and “quantity” were found to be overlooked by most frameworks, despite their growing impact for data practitioners as tools such as knowledge graphs become more common. Frameworks tailored to specific domains were also found to include fewer overall data quality dimensions but contained dimensions that were absent from more general frameworks, highlighting the need for a standardised approach that incorporates both established and emerging data quality dimensions. This work condenses information on commonly used and regulation-backed data quality frameworks, allowing practitioners to develop tools and applications to apply these frameworks that are compliant with standards and regulations. The bibliometric analysis from this review emphasises the importance of adopting a comprehensive quality framework to enhance governance, ensure regulatory compliance, and improve decision-making processes in data-rich environments.

  • Research Article
  • Cite Count Icon 6
  • 10.1177/0165551517748291
Big data to knowledge – Harnessing semiotic relationships of data quality and skills in genome curation work
  • Jan 11, 2018
  • Journal of Information Science
  • Hong Huang

This article aims to understand the views of genomic scientists with regard to the data quality assurances associated with semiotics and data–information–knowledge (DIK). The resulting communication of signs generated from genomic curation work, was found within different semantic levels of DIK that correlate specific data quality dimensions with their respective skills. Syntactic data quality dimensions were ranked the highest among all other semiotic data quality dimensions, which indicated that scientists spend great efforts for handling data wrangling activities in genome curation work. Semantic- and pragmatic-related sign communications were about meaningful interpretation, thus required additional adaptive and interpretative skills to deal with data quality issues. This expanded concept of ‘curation’ as sign/semiotic was not previously explored from the practical to the theoretical perspectives. The findings inform policy makers and practitioners to develop framework and cyberinfrastructure that facilitate the initiatives and advocacies of ‘Big Data to Knowledge’ by funding agencies. The findings from this study can also help plan data quality assurance policies and thus maximise the efficiency of genomic data management. Our results give strong support to the relevance of data quality skills communication for relationship with data quality assurance in genome curation activities.

  • Research Article
  • 10.1016/j.jjimei.2026.100407
The DaTUM framework: a multi-sector thematic analysis of data quality dimensions and their impacting factors
  • Jun 1, 2026
  • International Journal of Information Management Data Insights
  • Eleanor Smallwood + 3 more

• Multi-sector investigation into data quality dimensions and impacting factors. • A new theoretical framework to support organisational digitalisation or diversification. • A three-step process to apply the framework is presented to evaluate data quality and identify areas for improvement. • Thematic analysis of interviews with 31 practitioners from five sectors. • Four segments of data quality dimensions and five impacting factors are identified. The quality of data is central to decision making across all sectors. However, the many single-application studies, proposing hundreds of dimensions, make data quality assurance a daunting prospect. The diversification of industries is compounding this challenge, making a multi-sector classification essential. We propose a new classification for data quality – the DaTUM framework – which can act as a starting point for data quality assurance across diverse sectors. The framework was developed from reflective thematic analysis of 31 in-depth semi-structured interviews with academics and industry professionals from engineering, policy, economics, computer science, and psychology domains who generate, process, or analyse data. Eleven data quality dimensions and five factors that impact data quality were identified. The DaTUM framework, along with the operationalisation process, will simplify digitalisation and diversification efforts within organisations and support targeted data management and quality improvement strategies which are vital to achieving the true value of data.

  • Book Chapter
  • Cite Count Icon 8
  • 10.1016/b978-0-12-373717-5.00008-7
8 - Dimensions of Data Quality
  • Nov 18, 2010
  • The Practitioner's Guide to Data Quality Improvement
  • David Loshin

8 - Dimensions of Data Quality

  • Research Article
  • 10.47007/komp.v2i02.2190
PENGARUH KUALITAS DATA TERHADAP MANFAAT YANG DIPEROLEH DARI IMPLEMENTASI DATA WAREHOUSE
  • Jan 1, 2017
  • JIK: Jurnal Ilmu Komputer
  • Munawar Munawar

Given that the literature indicates that many benefits can be obtained from a DW (data warehouse), but that the threat of failure is also high, therefore more understanding of the quality dimensions that contribute to the success of a DW is required. This study address that need through an empirical investigation of success in data warehousing as measured by DQ (data quality) dimensions involvement in every stage of DW development. This study attempts to confirm DW benefits from reviewed literature in five case studies and three consultants. This study also identifies relationship amongst DQ dimensions in DW development stages and DW success. This study is helpful for DW practitioners, implementers and researchers for understanding of the challenges which DQ dimensions in DW stages are significant to establishing a successful DW development. Keywords: data warehouse, data quality, data warehouse benefits. Abstrak

  • Research Article
  • Cite Count Icon 57
  • 10.1016/j.im.2012.10.001
A multidimensional analysis of data quality for credit risk management: New insights and challenges
  • Nov 16, 2012
  • Information & Management
  • Helen-Tadesse Moges + 3 more

A multidimensional analysis of data quality for credit risk management: New insights and challenges

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 5
  • 10.1186/s13104-018-3161-8
Measuring management\u2019s perspective of data quality in Pakistan\u2019s Tuberculosis control programme: a test-based approach to identify data quality dimensions
  • Jan 16, 2018
  • BMC Research Notes
  • Syed Mustafa Ali + 5 more

BackgroundData quality is core theme of programme’s performance assessment and many organizations do not have any data quality improvement strategy, wherein data quality dimensions and data quality assessment framework are important constituents. As there is limited published research about the data quality specifics that are relevant to the context of Pakistan’s Tuberculosis control programme, this study aims at identifying the applicable data quality dimensions by using the ‘fitness-for-purpose’ perspective.ResultsForty-two respondents pooled a total of 473 years of professional experience, out of which 223 years (47%) were in TB control related programmes. Based on the responses against 11 practical cases, adopted from the routine recording and reporting system of Pakistan’s TB control programme (real identities of patient were masked), completeness, accuracy, consistency, vagueness, uniqueness and timeliness are the applicable data quality dimensions relevant to the programme’s context, i.e. work settings and field of practice.ConclusionBased on a ‘fitness-for-purpose’ approach to data quality, this study used a test-based approach to measure management’s perspective and identified data quality dimensions pertinent to the programme and country specific requirements. Implementation of a data quality improvement strategy and achieving enhanced data quality would greatly help organizations in promoting data use for informed decision making.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant