Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

A Novel Multi-layer Task-centric and Data Quality Framework for Autonomous Driving

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

The next-generation autonomous vehicles (AVs), embedded with frequent real-time decision-making, will rely heavily on massive multisource and multimodal data. In real-world settings, the data quality (DQ) of different sources and modalities varies due to unexpected environmental factors or sensor issues. However, both researchers and practitioners in the AV field overwhelmingly concentrate on models/algorithms while undervaluing the DQ. To improve AV’s functionality, efficiency, and trustworthiness, this paper proposes a novel task-centric and data quality vase framework, which consists of five layers: data, DQ, task, application, and goal. The proposed framework aims to map DQ with task requirements and performance goals. This paper opens up a range of critical but unexplored challenges at the intersection of DQ, task orchestration, and performance-oriented system development in AVs. It is expected to guide the AV community toward building more adaptive, explainable, and resilient AVs that respond intelligently to dynamic environments and heterogeneous data streams.

Similar Papers
  • Research Article
  • Cite Count Icon 39
  • 10.13063/2327-9214.1201
Using a Data Quality Framework to Clean Data Extracted from the Electronic Health Record: A Case Study
  • Jun 24, 2016
  • eGEMs
  • Oliwier Dziadkowiec + 4 more

Objectives:We examine the following: (1) the appropriateness of using a data quality (DQ) framework developed for relational databases as a data-cleaning tool for a data set extracted from two EPIC databases, and (2) the differences in statistical parameter estimates on a data set cleaned with the DQ framework and data set not cleaned with the DQ framework.Background:The use of data contained within electronic health records (EHRs) has the potential to open doors for a new wave of innovative research. Without adequate preparation of such large data sets for analysis, the results might be erroneous, which might affect clinical decision-making or the results of Comparative Effectives Research studies.Methods:Two emergency department (ED) data sets extracted from EPIC databases (adult ED and children ED) were used as examples for examining the five concepts of DQ based on a DQ assessment framework designed for EHR databases. The first data set contained 70,061 visits; and the second data set contained 2,815,550 visits. SPSS Syntax examples as well as step-by-step instructions of how to apply the five key DQ concepts these EHR database extracts are provided.Conclusions:SPSS Syntax to address each of the DQ concepts proposed by Kahn et al. (2012)1 was developed. The data set cleaned using Kahn’s framework yielded more accurate results than the data set cleaned without this framework. Future plans involve creating functions in R language for cleaning data extracted from the EHR as well as an R package that combines DQ checks with missing data analysis functions.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 7
  • 10.1051/rees/2022003
Analysing data quality frameworks and evaluating the statistical output of United Nations Sustainable Development Goals’ reports
  • Jan 1, 2022
  • Renewable Energy and Environmental Sustainability
  • Wajdi Al-Salim + 2 more

This paper evaluates the quality of the United Nations Sustainable Development Goals’ report for 2020, and devises a new data quality assessment framework based on analysing many data quality frameworks. Data in this paper is collected from the official UN SDG official website, and the national statistics offices of the UN countries. A weighted-score sum module is also being utilized to find the best data quality dimension. These dimensions are then used to create a new data quality framework. It is found that the UN SDGs used a data quality framework that is based on statistical output factors and ignores other quality factors and therefore the score for assessing this report is 56%. The perceived identified gaps include: countries are using different quality and assessment frameworks which cause inconsistency in data quality; data is outdated and incomplete; data is not available for many indicators and countries; cost and efficiency are not part of the UN SDG data quality framework; therefore weak data management is found. Areas for improvement include creating one comprehensive data quality framework for all countries will ensure the highest data quality.

  • Research Article
  • 10.55041/ijsrem8823
Data Quality Frameworks for Fraud Detection in Financial Reporting Pipelines
  • May 29, 2021
  • INTERNATIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT
  • Ravi Kiran Alluri

Abstract- Transparency, regulatory compliance, and trust in financial ecosystems depend on the integrity of financial reporting pipelines. Financial reporting fraud remains a widespread problem with serious repercussions for markets and stakeholders. Ensuring data quality at every stage is crucial for operational efficiency and successful fraud detection, as businesses depend increasingly on automated data pipelines for real-time financial reporting. With an emphasis on fraud detection capabilities, this paper investigates the development and application of comprehensive data quality frameworks suited to financial reporting pipelines. Post-hoc audits and manual checks are frequently the mainstays of traditional fraud detection methods, which are inadequate for managing large volumes of financial data at high speeds. The credibility of analytics and regulatory reporting can also be weakened by missing, erroneous, and inconsistent data in reporting systems, which can conceal fraudulent activity or produce false positives. Therefore, a strong data quality framework must incorporate domain-specific integrity constraints, anomaly detection logic, lineage tracing mechanisms, and enforce standard data validation rules. In line with anti-fraud goals, this study suggests a structured framework for data quality that combines rule-based, statistical, and metadata-driven validation procedures. Five essential pillars—completeness, accuracy, consistency, timeliness, and integrity—are incorporated into the framework presented in this paper. To identify questionable anomalies and deviations early, each pillar is mapped to particular validation mechanisms, including cross-ledger balancing, reference reconciliation, threshold-based monitoring, schema enforcement, and duplication checks. By incorporating these dimensions into ETL pipelines, organizations can proactively evaluate and score the quality of incoming data and flag records that might point to fraudulent manipulation, such as manipulated ledger entries, underreported liabilities, or revenue misstatements. The framework is integrated into a modular architecture that guarantees practical applicability with contemporary cloud-native data platforms and legacy systems. It uses tools like SQL-based integrity rules to identify transaction irregularities, Apache Atlas to track lineage and transformations, and Apache NiFi to orchestrate validation workflows. Financial controllers and compliance officers are also given access to real-time metrics and dashboards to monitor fraud risk indicators and data quality scores. This study uses a synthetic financial dataset enhanced with known fraud scenarios to assess the efficacy of the suggested framework. The analysis shows that while low-quality data introduces a lot of noise and lowers model reliability, high data quality scores are strongly correlated with fewer false positives in fraud detection models. The study also demonstrates how incorporating data quality validation early in the pipeline lifecycle enhances reporting output trust and speeds up the identification of fraudulent trends. This framework aids in the creation of safe, audit-ready, and compliant financial reporting pipelines by coordinating data quality assurance with fraud detection objectives. Additionally, it fills a significant void in managing financial data, where data quality is frequently viewed as an operational issue rather than a security or compliance requirement. The research's conclusions apply to financial institutions, regulators, auditors, and tech companies who want to improve financial reporting systems' resistance to fraud. This paper argues for a paradigm change in which data quality frameworks are essential to financial reporting fraud detection rather than being merely incidental. Adopting such frameworks will be crucial to creating robust, transparent, and reliable financial ecosystems as the financial sector accelerates digital transformation. Future research could build on this work by incorporating machine learning methods into the data quality scoring process, enabling even more automation and accuracy in fraud detection workflows. Keywords: Data quality, fraud detection, financial reporting pipelines, ETL validation, financial integrity, data lineage, anomaly detection, audit compliance, data governance, metadata-driven validation.

  • Research Article
  • 10.1182/blood-2025-4443
Data validation and quality framework for building a european multimodal real-world dataset for clinical outcome research in sickle cell disease
  • Nov 3, 2025
  • Blood
  • Sara Reidel + 19 more

Data validation and quality framework for building a european multimodal real-world dataset for clinical outcome research in sickle cell disease

  • Research Article
  • Cite Count Icon 2
  • 10.14329/apjis.2020.30.2.420
Factors Influencing the Knowledge Adoption of Mobile Game Developers in Online Communities : Focusing on the HSM and Data Quality Framework
  • Jun 30, 2020
  • Asia Pacific Journal of Information Systems
  • Jong-Won Park + 2 more

Recently, with the advance of the wireless Internet access via mobile devices, a myriad of game development companies have forayed into the mobile game market, leading to intense competition. To survive in this fierce competition, mobile game developers often try to get a grasp of the rapidly changing needs of their customers by operating their own official communities where game users freely leave their requests, suggestions, and ideas relevant to focal games. Based on the heuristic-systematic model (HSM) and the data quality (DQ) framework, this study derives key content, non-content, and hybrid cues that can be utilized when game developers accept suggested postings in these online communities. The results of hierarchical multiple regression analysis show that relevancy, timeliness, amount of writing, and the number of comments are positively associated with mobile game developers’ knowledge adoption. In addition, title attractiveness mitigates the relationship between amount of writing/the number of comments and knowledge adoption.

  • Research Article
  • Cite Count Icon 5
  • 10.3233/sji-230125
Evaluating data quality for blended data using a data quality framework.
  • Feb 1, 2024
  • Statistical journal of the IAOS
  • Jennifer D Parker + 5 more

In 2020 the U.S. Federal Committee on Statistical Methodology (FCSM) released "A Framework for Data Quality", organized by 11 dimensions of data quality grouped among three domains of quality (utility, objectivity, integrity). This paper addresses the use of the FCSM Framework for data quality assessments of blended data. The FCSM Framework applies to all types of data, however best practices for implementation have not been documented. We applied the FCSM Framework for three health-research related case studies. For each case study, assessments of data quality dimensions were performed to identify threats to quality, possible mitigations of those threats, and trade-offs among them. From these assessments the authors concluded: 1) data quality assessments are more complex in practice than anticipated and expert guidance and documentation are important; 2) each dimension may not be equally important for different data uses; 3) data quality assessments can be subjective and having a quantitative tool could help explain the results, however, quantitative assessments may be closely tied to the intended use of the dataset; 4) there are common trade-offs and mitigations for some threats to quality among dimensions. This paper is one of the first to apply the FCSM Framework to specific use-cases and illustrates a process for similar data uses.

  • Research Article
  • Cite Count Icon 13
  • 10.3390/bdcc9040093
A Comparison of Data Quality Frameworks: A Review
  • Apr 9, 2025
  • Big Data and Cognitive Computing
  • Russell Miller + 3 more

This study reviews various data quality frameworks that have some form of regulatory backing. The aim is to identify how these frameworks define, measure, and apply data quality dimensions. This review identified generalisable frameworks, such as TDQM, ISO 8000, and ISO 25012, and specialised frameworks, such as IMF’s DQAF, BCBS 239, WHO’s DQA, and ALCOA+. A standardised data quality model was employed to map the dimensions of the data from each framework to a common vocabulary. This mapping enabled a gap analysis that highlights the presence or absence of specific data quality dimensions across the examined frameworks. The analysis revealed that core data quality dimensions such as “accuracy”, “completeness”, “consistency”, and “timeliness” are equally and well represented across all frameworks. In contrast, dimensions such as “semantics” and “quantity” were found to be overlooked by most frameworks, despite their growing impact for data practitioners as tools such as knowledge graphs become more common. Frameworks tailored to specific domains were also found to include fewer overall data quality dimensions but contained dimensions that were absent from more general frameworks, highlighting the need for a standardised approach that incorporates both established and emerging data quality dimensions. This work condenses information on commonly used and regulation-backed data quality frameworks, allowing practitioners to develop tools and applications to apply these frameworks that are compliant with standards and regulations. The bibliometric analysis from this review emphasises the importance of adopting a comprehensive quality framework to enhance governance, ensure regulatory compliance, and improve decision-making processes in data-rich environments.

  • Research Article
  • 10.1145/3768344
Supporting Timing-related Metrics for Autonomous Driving Frameworks in CyberRT
  • Nov 11, 2025
  • ACM Transactions on Design Automation of Electronic Systems
  • Miguel Alcon + 3 more

The provision of increasingly advanced autonomous software functionalities builds on cutting-edge autonomous driving frameworks to enable modular interactions among multiple software components. This approach helps to support functional cause-effect chains from multiple sensors to actuators. The complexity of the (software) component interactions makes it more difficult to ascertain the correctness of the timing behavior of the system. This is so because traditional timing-related metrics like worst-case execution and worst-case response time do not capture the inter-dependency in cause-effect chains between the input sampling time and the time at which computation based on those inputs is performed. Complementary timing-related metrics, such as maximum reaction time and maximum data age have been considered to capture timing requirements, typically with an end-to-end scope, in cause-effect chains. These metrics have been formalized and demonstrated in ROS2-based automotive and autonomous driving setups [ 44 , 46 ]. However, the formalization of those metrics, which is necessary for deriving analytical lower and upper bounds and monitoring them at run-time, largely depends on the execution model and semantics offered by the run-time. Any concrete application of those metrics need to be tailored and adapted to the system at hand. Apollo auto is a popular, industrial-quality, open-source autonomous driving framework that is seeing increasing adoption both for industrial and academic projects. Apollo builds on CyberRT , an ad-hoc run-time that is similar in mechanism and intent to ROS2 but differentiates from it with respect to execution model and supported semantics. In contrast to ROS2, CyberRT is highly specialized to support the Apollo AD framework, is neither extensively documented or thoroughly analysed in the literature, especially in relation to execution model and instantiation of timing-related metrics. In this work, for the first time, we provide an insightful analysis and discussion on CyberRT execution model and semantics, starting from its raw and non-extensively documented codebase. Based on the identified semantics, we elaborate a formalization of timing-related metrics on CyberRT , across different granularity scopes, namely end-to-end and node levels. In particular, we develop on the importance of node-level timing properties to intercept any latent timing misbehavior before it is too late, and it severely impacts end-to-end execution. We provide a concrete mapping of a comprehensive set of timing-related metrics to the CyberRT execution model, both at end-to-end and node level, and develop a monitoring library that allows to intercept them on the specific software stack. We exploit the proposed library on a set of Apollo autonomous driving scenarios to demonstrate its effectiveness in monitoring the considered timing metrics and to promptly intercept a subtle timing misbehavior beyond end-to-end execution scope in a representative autonomous driving stack.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 15
  • 10.5334/egems.298
Improving a Secondary Use Health Data Warehouse: Proposing a Multi-Level Data Quality Framework
  • Aug 2, 2019
  • eGEMs
  • Sandra Henley-Smith + 2 more

Background:Data quality frameworks within information technology and recently within health care have evolved considerably since their inception. When assessing data quality for secondary uses, an area not yet addressed adequately in these frameworks is the context of the intended use of the data.Methods:After review of literature to identify relevant research, an existing data quality framework was refined and expanded to encompass the contextual requirements not present.Results:The result is a two-level framework to address the need to maintain the intrinsic value of the data, as well as the need to indicate whether the data will be able to provide the basis for answers in specific areas of interest or questions.Discussion:Data quality frameworks have always been one dimensional, requiring the implementers of these frameworks to fit the requirements of the data’s use around how the framework is designed to function. Our work has systematically addressed the shortcomings of existing frameworks, through the application of concepts synthesized from the literature to the naturalistic setting of data quality management in an actual health data warehouse.Conclusion:Secondary use of health data relies on contextualized data quality management. Our work is innovative in showing how to apply context around data quality characteristics and how to develop a second level data quality framework, so as to ensure that quality and context are maintained and addressed throughout the health data quality assessment process.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 123
  • 10.1186/s12874-021-01252-7
Facilitating harmonized data quality assessments. A data quality framework for observational health research data collections with software implementations in R
  • Apr 2, 2021
  • BMC Medical Research Methodology
  • Carsten Oliver Schmidt + 9 more

BackgroundNo standards exist for the handling and reporting of data quality in health research. This work introduces a data quality framework for observational health research data collections with supporting software implementations to facilitate harmonized data quality assessments.MethodsDevelopments were guided by the evaluation of an existing data quality framework and literature reviews. Functions for the computation of data quality indicators were written in R. The concept and implementations are illustrated based on data from the population-based Study of Health in Pomerania (SHIP).ResultsThe data quality framework comprises 34 data quality indicators. These target four aspects of data quality: compliance with pre-specified structural and technical requirements (integrity); presence of data values (completeness); inadmissible or uncertain data values and contradictions (consistency); unexpected distributions and associations (accuracy). R functions calculate data quality metrics based on the provided study data and metadata and R Markdown reports are generated. Guidance on the concept and tools is available through a dedicated website.ConclusionsThe presented data quality framework is the first of its kind for observational health research data collections that links a formal concept to implementations in R. The framework and tools facilitate harmonized data quality assessments in pursue of transparent and reproducible research. Application scenarios comprise data quality monitoring while a study is carried out as well as performing an initial data analysis before starting substantive scientific analyses but the developments are also of relevance beyond research.

  • Abstract
  • 10.1016/j.jval.2023.03.2125
RWD96 Implementation of a Real-World Data Quality Framework in a Nationwide Oncology Electronic Health Record-Derived Database
  • Jun 1, 2023
  • Value in Health
  • E Castellanos + 2 more

RWD96 Implementation of a Real-World Data Quality Framework in a Nationwide Oncology Electronic Health Record-Derived Database

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 9
  • 10.1177/08944393241245395
Assessing Data Quality in the Age of Digital Social Research: A Systematic Review
  • Apr 27, 2024
  • Social Science Computer Review
  • Jessica Daikeler + 9 more

While survey data has long been the focus of quantitative social science analyses, observational and content data, although long-established, are gaining renewed attention; especially when this type of data is obtained by and for observing digital content and behavior. Today, digital technologies allow social scientists to track “everyday behavior” and to extract opinions from public discussions on online platforms. These new types of digital traces of human behavior, together with computational methods for analyzing them, have opened new avenues for analyzing, understanding, and addressing social science research questions. However, even the most innovative and extensive amounts of data are hollow if they are not of high quality. But what does data quality mean for modern social science data? To investigate this rather abstract question the present study focuses on four objectives. First, we provide researchers with a decision tree to identify appropriate data quality frameworks for a given use case. Second, we determine which data types and quality dimensions are already addressed in the existing frameworks. Third, we identify gaps with respect to different data types and data quality dimensions within the existing frameworks which need to be filled. And fourth, we provide a detailed literature overview for the intrinsic and extrinsic perspectives on data quality. By conducting a systematic literature review based on text mining methods, we identified and reviewed 58 data quality frameworks. In our decision tree, the three categories, namely, data type, the perspective it takes, and its level of granularity, help researchers to find appropriate data quality frameworks. We, furthermore, discovered gaps in the available frameworks with respect to visual and especially linked data and point out in our review that even famous frameworks might miss important aspects. The article ends with a critical discussion of the current state of the literature and potential future research avenues.

  • Research Article
  • Cite Count Icon 5
  • 10.1109/jsyst.2020.2985343
Framework for Integral Data Quality and Security Evaluation in Smartphones
  • Jun 1, 2021
  • IEEE Systems Journal
  • Igor Khokhlov + 2 more

Data quality (DQ) concept should play an important role in decision-making and engineering systems. Underestimation of DQ may lead to resource waste, wrong conclusions, or inefficient decisions. Unfortunately, current approaches to DQ incorporating into data management systems are limited to particular applications. This problem is aggravated by the DQ inequality of data sources. This is especially critical in mobile crowd-sensing applications where data may come from unverified data contributors using the smartphones and other mobile devices. To facilitate the expansion of DQ evaluation to a wider spectrum of applications, this article presents a framework for integral DQ and security evaluation in Android-based smartphones. The developed framework provides support for selecting the DQ metrics and implementing their calculus by integrating diverse sensor DQ and security metrics. We present multiple calculi for DQ and security evaluation such as hierarchical fuzzy rules expert system, neural networks, and algebraic functions. Case studies that demonstrate the framework's performance in addressing real-life tasks are presented and the achieved results are analyzed. The implementation results validate the framework's capability of performing comprehensive DQ evaluations.

  • Abstract
  • Cite Count Icon 1
  • 10.1093/eurpub/ckac129.498
A Data Quality Framework for the European Health Data Space for secondary use
  • Oct 21, 2022
  • The European Journal of Public Health
  • E Bernal-Delgado + 4 more

A cornerstone in the development of the European Health Data Space for secondary use of data (EHDS2) is the design, implementation and assessment of a Data Quality Framework (DQF). Consistently, the Joint Action TEHDAS has a dedicated work program where, learning from others’ experiences across Europe and abroad, the work package is building the concepts and methods for such a DQF. The scope of this work program is to provide recommendation to the Member States and the European Commission on the concept of DQF to foster, where (institutions) the DQF should be implemented, when in data life cycle, how should be implemented and by whom. In terms of the concept, the DQF raises the importance of quality assurance procedures at data processor level and the level of quality of the data collections in terms of reliability, relevance, timeliness, coherence, coverage and completeness. When it comes to when along the data life cycle, DQF is expected operate when data needs harmonization at data processor level (ie, the effective application of interoperability standards), in the publication of the data sources (ie providing users knowledge on the provenance of data and the content of data source); or, when data sources have to be integrated and sensitive data pseudonymized (ie, the quality of the linkage and losses after pseudonymisation). Finally, when it comes to the methodology, TEHDAS suggests a three-fold approach - some quality measures in the DQF could be translated into legislation (eg, the requirement of regular auditing for a data processor to be a trusted party in the EDHS2); some could be kept as good-practices (eg, recommendation of archival procedures when a research project finalizes); and, under the assumption of continuous data quality improvement, an assessment, benchmarking and promotion methodology (eg, a grading system at data processor level).

  • Research Article
  • 10.1038/s41598-025-27787-z
RAUM-GANs: a multi-layer GAN-enhanced framework for accurate multiple sclerosis lesion segmentation in MRI
  • Dec 16, 2025
  • Scientific Reports
  • Ahmed Alsayat + 10 more

Multiple sclerosis (MS) is a chronic autoimmune disease characterized by inflammatory brain lesions, making MRI-based lesion segmentation challenging due to noise, missing data, and limited availability of high-quality labeled images. This paper presents RAUM-GANs, a multi-layer deep learning framework designed to address these challenges and enhance segmentation accuracy. The preprocessing stage comprises three layers: (1) noise reduction using a modified Denoising GAN (DGAN-Net), achieving peak signal-to-noise ratio (PSNR) values up to 42.21 dB across varying noise levels; (2) missing data imputation through advanced GAN-based methods, ensuring clinically reliable reconstruction of incomplete MRI scans; and (3) dataset expansion via a Multi-level Identity GAN (MGAN), which incorporates an identity block to prevent mode collapse, an 8-connected pixel constraint to maintain spatial coherence, and a softened discriminator output to mitigate vanishing gradients. For segmentation, a Residual Attention U-Net (RAU-Net) with identity mapping is employed, yielding precise detection and delineation of MS lesions. Extensive evaluation on the MICCAI MSSEG-2 dataset demonstrates that RAUM-GANs outperform four state-of-the-art methods, achieving a Dice score of 96.6%, Fréchet Inception Distance (FID) of 43.13, and Inception Score (IS) of 14.03. The results highlight the framework’s ability to generate high-quality synthetic MRI data, improve robustness against noise and incomplete information, and deliver superior lesion segmentation performance. RAUM-GANs provides a comprehensive, scalable solution for MS lesion analysis, with potential applicability to other medical imaging domains where data quality and scarcity remain significant barriers.

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant