Comparison of Imputation Methods for Missing Values in Air Pollution Data: Case Study on Sydney Air Quality Index

W M L K N Wijesekara,Liwan Liyanage

doi:10.1007/978-3-030-39442-4_20

Abstract

Missing values in air quality data may lead to a substantial amount of bias and inefficiency in modeling. In this paper, we discuss six methods for dealing with missing values in univariate time series and compare their performances. The methods we discuss here are Mean Imputation, Spline Interpolation, Simple Moving Average, Exponentially Weighted Moving Average, Kalman Smoothing on Structural Time Series Models and Kalman Smoothing on Autoregressive Integrated Moving Average (ARIMA) models. The performances of these methods were compared using three different performance measures; Mean Squared Error, Coefficient of Determination and the Index of Agreement. Kalman Smoothing on Structural Time Series method is the best method among the methods considered, for imputing missing values in the context of air quality data under Missing Completely at Random (MCAR) mechanism. Kalman Smoothing on ARIMA, and Exponentially Weighted Moving Average methods also perform considerably well. Performance of Spline Interpolation decreases drastically with increased percentage of missing values. Mean Imputation performs reasonably well for smaller percentage of missing values; however, all the other methods outperform Mean Imputation regardless the number of missing values.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Comparison of Imputation Methods for Missing Values in Air Pollution Data: Case Study on Sydney Air Quality Index

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Air quality data pre-processing: A novel algorithm to impute missing values in univariate time series
Lakmini Wijesekara ... Liwan Liyanage
-
Lakmini Wijesekara, et. al.Lakmini Wijesekara ... Liwan Liyanage
01 Nov 2021
01 Nov 2021

Mind the Large Gap: Novel Algorithm Using Seasonal Decomposition and Elastic Net Regression to Impute Large Intervals of Missing Data in Air Quality Data
Lakmini Wijesekara ... Liwan Liyanage
Atmosphere | VOL. 14
Lakmini Wijesekara, et. al.Lakmini Wijesekara ... Liwan Liyanage
10 Feb 2023
Atmosphere | VOL. 14

Making sense of antimicrobial use and resistance surveillance data: application of ARIMA and transfer function models
D.L Monnet ... N Gonzalo
Clinical Microbiology and Infection | VOL. 7
D.L Monnet, et. al.D.L Monnet ... N Gonzalo
01 Jan 2001
Clinical Microbiology and Infection | VOL. 7

Characteristics of Machine Learning-based Univariate Time Series Imputation Method
Dini Ramadhani ... Erfiani Erfiani
JUITA: Jurnal Informatika | VOL. 12
Dini Ramadhani, et. al.Dini Ramadhani ... Erfiani Erfiani
07 Nov 2024
JUITA: Jurnal Informatika | VOL. 12

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Comparison of Imputation Methods for Missing Values in Air Pollution Data: Case Study on Sydney Air Quality Index

Abstract

Talk to us

Similar Papers