Efficient Model-Free Subsampling Method for Massive Data

Zheng Zhou,Zebin Yang,Aijun Zhang,Yongdao Zhou

doi:10.1080/00401706.2023.2271091

Abstract

Subsampling plays a crucial role in tackling problems associated with the storage and statistical learning of massive datasets. However, most existing subsampling methods are model-based, which means their performances can drop significantly when the underlying model is misspecified. Such an issue calls for model-free subsampling methods that are robust under diverse model specifications. Recently, several model-free subsampling methods have been developed. However, the computing time of these methods grows explosively with the sample size, making them impractical for handling massive data. In this article, an efficient model-free subsampling method is proposed, which segments the original data into some regular data blocks and obtains subsamples from each data block by the data-driven subsampling method. Compared with existing model-free subsampling methods, the proposed method has a significant speed advantage and performs more robustly for datasets with complex underlying distributions. As demonstrated in simulation experiments, the proposed method is an order of magnitude faster than other commonly used model-free subsampling methods when the sample size of the original dataset reaches the order of 107. Moreover, simulation experiments and case studies show that the proposed method is more robust than other model-free subsampling methods under diverse model specifications and subsample sizes.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Efficient Model-Free Subsampling Method for Massive Data

Abstract

Talk to us

Similar Papers

More From: Technometrics

Lead the way for us

Journal: Technometrics	Publication Date: Nov 18, 2023
Citations: 2

Similar Papers

Subsample and half-sample methods
Gutti Jogesh Babu
Annals of the Institute of Statistical Mathematics | VOL. 44
Gutti Jogesh BabuGutti Jogesh Babu
01 Dec 1992
Annals of the Institute of Statistical Mathematics | VOL. 44

An Optimal Transport Approach for Selecting a Representative Subsample with Application in Efficient Kernel Density Estimation
Jingyi Zhang ... Ping Ma
Journal of Computational and Graphical Statistics | VOL. 32
Jingyi Zhang, et. al.Jingyi Zhang ... Ping Ma
03 Jun 2022
Journal of Computational and Graphical Statistics | VOL. 32

Automatic background correction method for laser-induced breakdown spectroscopy
Hao Chen ... Wenli Zhang
Spectrochimica Acta Part B: Atomic Spectroscopy | VOL. 208
Hao Chen, et. al.Hao Chen ... Wenli Zhang
05 Aug 2023
Spectrochimica Acta Part B: Atomic Spectroscopy | VOL. 208

Fuzzy neuron model-free control for continuous steel casting processes
Wang Ning
-
Wang Ning Wang Ning
28 Jun 2000
28 Jun 2000

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Efficient Model-Free Subsampling Method for Massive Data

Abstract

Talk to us

Similar Papers

More From: Technometrics