Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation

Yicheng Yang,In Ho Cho,Jae Kwang Kim,Yonghyun Kwon

doi:10.1109/tkde.2023.3249567

Yicheng Yang, In Ho Cho + Show 2 more

Open Access

https://doi.org/10.1109/tkde.2023.3249567

Copy DOI

Abstract

Parallel fractional hot-deck imputation (P-FHDI [1]) is a general-purpose, assumption-free tool for handling item nonresponse in big incomplete data by combining the theory of FHDI and parallel computing. FHDI cures multivariate missing data by filling each missing unit with multiple observed values (thus, hot-deck) without resorting to distributional assumptions. P-FHDI can tackle big incomplete data with millions of instances (big- <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$n$</tex-math></inline-formula> ) or 10,000 variables (big- <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$p$</tex-math></inline-formula> ). However, handling ultra incomplete data (i.e., concurrently big- <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$n$</tex-math></inline-formula> and big- <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$p$</tex-math></inline-formula> ) with tremendous instances and high dimensionality has posed problems to P-FHDI due to excessive memory requirement and execution time. To tackle the aforementioned challenges, we propose the ultra data-oriented P-FHDI (named UP-FHDI) capable of curing ultra incomplete data. In addition to the parallel Jackknife method, this paper enables a computationally efficient ultra data-oriented variance estimation using parallel linearization techniques. Results confirm that UP-FHDI can tackle an ultra dataset with one million instances and 10,000 variables. This paper illustrates the special parallel algorithms of UP-FHDI and confirms its positive impact on the subsequent deep learning performance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering

Lead the way for us

Journal: IEEE Transactions on Knowledge and Data Engineering	Publication Date: Sep 1, 2023
License type: CC BY 4.0

Similar Papers

Second-order variance estimation in poststratified two-stage sampling
Kyongryun Kim ... Suojin Wang
Journal of Statistical Planning and Inference | VOL. 139
Kyongryun Kim, et. al.Kyongryun Kim ... Suojin Wang
21 Nov 2008
Journal of Statistical Planning and Inference | VOL. 139

Variance Estimation for the Regression Estimator in Two-Phase Sampling
R R Sitter
Journal of the American Statistical Association | VOL. 92
R R SitterR R Sitter
01 Jun 1997
Journal of the American Statistical Association | VOL. 92

A new replicate variance estimator for unequal probability sampling without replacement
Emilio L Escobar ... Yves G Berger
Canadian Journal of Statistics | VOL. 41
Emilio L Escobar, et. al.Emilio L Escobar ... Yves G Berger
19 Jul 2013
Canadian Journal of Statistics | VOL. 41

Variance Estimation for the Regression Estimator in Two-Phase Sampling
R R Sitter
Journal of the American Statistical Association | VOL. 92
R R SitterR R Sitter
01 Jun 1997
Journal of the American Statistical Association | VOL. 92

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Ultra Data-Oriented Parallel Fractional Hot-Deck Imputation With Efficient Linearized Variance Estimation

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Knowledge and Data Engineering