An Empirical Evaluation of Feature Selection Stability and Classification Accuracy

Mustafa Büyükkeçeci̇,Mehmet Cudi Okur

doi:10.35378/gujs.998964

Abstract

The performance of inductive learners can be negatively affected by high-dimensional datasets. To address this issue, feature selection methods are used. Selecting relevant features and reducing data dimensions is essential for having accurate machine learning models. Stability is an important criterion in feature selection. Stable feature selection algorithms maintain their feature preferences even when small variations exist in the training set. Studies have emphasized the importance of stable feature selection, particularly in cases where the number of samples is small and the dimensionality is high. In this study, we evaluated the relationship between stability measures, as well as, feature selection stability and classification accuracy, using the Pearson Product-Moment Correlation Coefficient. We conducted an extensive series of experiments using five filter and two wrapper feature selection methods, three classifiers for subset and classification performance evaluation, and eight real-world datasets taken from two different data repositories. We measured the stability of feature selection methods using a total of twelve stability metrics. Based on the results of correlation analyses, we have found that there is a lack of substantial evidence supporting a linear relationship between feature selection stability and classification accuracy. However, a strong positive correlation has been observed among several stability metrics.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

An Empirical Evaluation of Feature Selection Stability and Classification Accuracy

Abstract

Talk to us

Similar Papers

More From: Gazi University Journal of Science

Lead the way for us

Similar Papers

Stability of Feature Selection Methods: A Study of Metrics Across Different Gene Expression Datasets
Zahra Mungloo-Dilmohamud ... Carlos Peña-Reyes
-
Zahra Mungloo-Dilmohamud, et. al.Zahra Mungloo-Dilmohamud ... Carlos Peña-Reyes
01 Jan 2020
01 Jan 2020

Feature Selection and Feature Stability Measurement Method for High-Dimensional Small Sample Data Based on Big Data Technology.
Chengyuan Huang
Computational Intelligence and Neuroscience | VOL. 2021
Chengyuan HuangChengyuan Huang
01 Jan 2020
Computational Intelligence and Neuroscience | VOL. 2021

On the Stability of Feature Selection Methods in Software Quality Prediction: An Empirical Investigation
Huanjing Wang ... Taghi M Khoshgoftaar
International Journal of Software Engineering and Knowledge Engineering | VOL. 25
Huanjing Wang, et. al.Huanjing Wang ... Taghi M Khoshgoftaar
01 Nov 2015
International Journal of Software Engineering and Knowledge Engineering | VOL. 25

Stability of Three Forms of Feature Selection Methods on Software Engineering Data
Huanjing Wang ... Amri Napolitano
-
Huanjing Wang, et. al.Huanjing Wang ... Amri Napolitano
01 Jul 2015
01 Jul 2015

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

An Empirical Evaluation of Feature Selection Stability and Classification Accuracy

Abstract

Talk to us

Similar Papers

More From: Gazi University Journal of Science