An Assessment of Open Data Sets Completeness

Abdulrazzak Ali,Amelia R,Siti A,Nurul A

doi:10.14569/ijacsa.2019.0100672

Abstract

The rapid growth of open data sources is driven by free-of-charge contents and ease of accessibility. While it is convenient for public data consumers to use data sets extracted from open data sources, the decision to use these data sets should be based on data sets’ quality. Several data quality dimensions such as completeness, accuracy, and timeliness are common requirements to make data fit for use. More importantly, in many cases, high-quality data sets are desirable in ensuring reliable outcomes of reports and analytics. Even though many open data sources provide data quality guidelines, the responsibility to ensure data of high quality requires commitment from data contributors. In this paper, an initial investigation on the quality of open data sets in terms of completeness dimension was con-ducted. In particular, the results of the missing values in 20 open data sets measurement were extracted from the open data sources. The analysis covered all the missing values representations which are not limited to nulls or blank spaces. The results exhibited a range of missing values ratios that indicated the level of the data sets completeness. The limited coverage of this analysis does not hinder understanding of the current level of data completeness of open data sets. The findings may motivate open data providers to design initiatives that will empower data quality policy and guidelines for data contributors. In addition, this analysis may assist public data users to decide on the acceptability of open data sets by applying the simple methods proposed in this paper or performing data cleaning actions to improve the completeness of the data sets concerned.

Highlights

Data completeness is an essential dimension in data quality like accuracy and timeliness
We presented the results of assessing data completeness problem in open data sets
The assessment results involving twenty open data sets show varying missing values ratios that perhaps can be explained by the nature of the data set, data collection policy and enforcement

Summary

INTRODUCTION

Data completeness is an essential dimension in data quality like accuracy and timeliness. The first case represents the total loss of information where the attributes’ values for the whole record are missing. Assume that we have a simple data set which is supposed to have ten records of students’ information. All records are uniquely identified by an identification attribute, StudentId. A missing record can be represented by the absence of the student’s record with id ‘B14’ from the data set.

BACKGROUND

Reasons of Missing Data

Missing Values Representation

Methods for Handling Missing Values

METHODOLOGY

Findings

CONCLUSIONS

Full Text

Published version (

Free)

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An Assessment of Open Data Sets Completeness

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications

Lead the way for us

Journal: International Journal of Advanced Computer Science and Applications	Publication Date: Jan 1, 2019
License type: cc-by

Similar Papers

Open Data Interface (ODI) for secondary school education
Mubashrah Saddiqa ... Jens Myrup Pedersen
Computers & Education | VOL. 174
Mubashrah Saddiqa, et. al.Mubashrah Saddiqa ... Jens Myrup Pedersen
11 Aug 2021
Computers & Education | VOL. 174

The Analysis of Open Source Software and Data for Establishment of GIS Services Throughout the Network in a Mapping Organization at National or International Level

-

01 Jan 2014
01 Jan 2014

Drivers and barriers towards re-using open government data (OGD): a case study of open data initiative in Oman
Stuti Saxena
foresight | VOL. 20
Stuti SaxenaStuti Saxena
09 Apr 2018
foresight | VOL. 20

Open Data Quality Evaluation: A Comparative Analysis of Open Data in Latvia
Anastasija Nikiforova
Baltic Journal of Modern Computing | VOL. 6
Anastasija NikiforovaAnastasija Nikiforova
01 Jan 2018
Baltic Journal of Modern Computing | VOL. 6

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An Assessment of Open Data Sets Completeness

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications