A Combined Feature Screening Approach of Random Forest and Filterbased Methods for Ultra-high Dimensional Data

Lifeng Zhou,Hong Wang

doi:10.2174/1574893617666220221120618

Abstract

Background: Various feature (variable) screening approaches have been proposed in the past decade to mitigate the impact of ultra-high dimensionality in classification and regression problems, including filter based methods such as sure independence screening, and wrapper based methods such as random forest. However, the former type of methods rely heavily on strong modelling assumptions while the latter ones requires an adequate sample size to make the data speak for themselves. These requirements can seldom be met in biochemical studies in cases where we have only access to ultra-high dimensional data with a complex structure and a small number of observations. Objective: In this research, we want to investigate the possibility of combining both filter based screening methods and random forest based screening methods in the regression context. Method: We have combined four state-of-art filter approaches, namely, sure independence screening (SIS), robust rank correlation based screening (RRCS), high dimensional ordinary least squares projection (HOLP) and a model free sure independence screening procedure based on the distance correlation (DCSIS) from the statistical community with a random forest based Boruta screening method from the machine learning community for regression problems. Result: Among all the combined methods, RF-DCSIS performs better than the other methods in terms of screening accuracy and prediction capability on the simulated scenarios and real benchmark datasets. Conclusion: By empirical study from both extensive simulation and real data, we have shown that both filter based screening and random forest based screening have their pros and cons, while a combination of both may lead to a better feature screening result and prediction capability.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Combined Feature Screening Approach of Random Forest and Filterbased Methods for Ultra-high Dimensional Data

Abstract

Talk to us

Similar Papers

More From: Current Bioinformatics

Lead the way for us

Journal: Current Bioinformatics	Publication Date: May 1, 2022
Citations: 12

Similar Papers

Combined performance of screening and variable selection methods in ultra-high dimensional data in predicting time-to-event outcomes
Lira Pi ... Susan Halabi
Diagnostic and Prognostic Research | VOL. 2
Lira Pi, et. al.Lira Pi ... Susan Halabi
26 Sep 2018
Diagnostic and Prognostic Research | VOL. 2

A Novel Feature Selection Method for Ultra High Dimensional Survival Data
Nahid Salma ... Majid Khan Majahar Ali
Malaysian Journal of Fundamental and Applied Sciences | VOL. 20
Nahid Salma, et. al.Nahid Salma ... Majid Khan Majahar Ali
15 Oct 2024
Malaysian Journal of Fundamental and Applied Sciences | VOL. 20

Feature Screening via Distance Correlation Learning
Runze Li ... Liping Zhu
Journal of the American Statistical Association | VOL. 107
Runze Li, et. al.Runze Li ... Liping Zhu
01 Jun 2012
Journal of the American Statistical Association | VOL. 107

Robust sure independence screening for ultrahigh dimensional non-normal data
Wei Zhong
Acta Mathematica Sinica, English Series | VOL. 30
Wei ZhongWei Zhong
15 Oct 2014
Acta Mathematica Sinica, English Series | VOL. 30

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Combined Feature Screening Approach of Random Forest and Filterbased Methods for Ultra-high Dimensional Data

Abstract

Talk to us

Similar Papers

More From: Current Bioinformatics