Handling high-dimensional data with missing values by modern machine learning techniques

Sixia Chen,Chao Xu

doi:10.1080/02664763.2022.2068514

Abstract

High-dimensional data have been regarded as one of the most important types of big data in practice. It happens frequently in practice including genetic study, financial study, and geographical study. Missing data in high dimensional data analysis should be handled properly to reduce nonresponse bias. We discuss some modern machine learning techniques including penalized regression approaches, tree-based approaches, and deep learning (DL) for handling missing data with high dimensionality. Specifically, our proposed methods can be used for estimating general parameters of interest including population means and percentiles with imputation-based estimators, propensity score estimators, and doubly robust estimators. We compare those methods through some limited simulation studies and a real application. Both simulation studies and real application show the benefits of DL and XGboost approaches compared with other methods in terms of balancing bias and variance.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Handling high-dimensional data with missing values by modern machine learning techniques

Abstract

Talk to us

Similar Papers

More From: Journal of Applied Statistics

Lead the way for us

Journal: Journal of Applied Statistics	Publication Date: May 3, 2022
Citations: 4

Similar Papers

Preface
S Ejaz Ahmed
Applied Stochastic Models in Business and Industry | VOL. 35
S Ejaz AhmedS Ejaz Ahmed
01 Mar 2019
Applied Stochastic Models in Business and Industry | VOL. 35

High-dimensional data: p >> n in mathematical statistics and bio-medical applications
Sara A Van De Geer ... Hans C Van Houwelingen
Bernoulli | VOL. 10
Sara A Van De Geer, et. al.Sara A Van De Geer ... Hans C Van Houwelingen
01 Dec 2004
Bernoulli | VOL. 10

Application of non parametric Bayesian methods in high dimensional data
Yunqing Xia
Journal of Computational Methods in Sciences and Engineering | VOL. 24
Yunqing XiaYunqing Xia
10 May 2024
Journal of Computational Methods in Sciences and Engineering | VOL. 24

Analysis of Patient Groups and Immunization Results Based on Subspace Clustering
Michael Hund ... Ljiljana Majnaric
-
Michael Hund, et. al.Michael Hund ... Ljiljana Majnaric
01 Jan 2015
01 Jan 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Handling high-dimensional data with missing values by modern machine learning techniques

Abstract

Talk to us

Similar Papers

More From: Journal of Applied Statistics