Comparison of Feature Selection and Feature Extraction Role in Dimensionality Reduction of Big Data

Haidar Khalid Malik Haidar Khalid Malik,Nashaat Jasim Al-Anber Nashaat Jasim Al-Anber

doi:10.51173/jt.v5i1.1027

Haidar Khalid Malik Haidar Khalid Malik, Nashaat Jasim Al-Anber Nashaat Jasim Al-Anber

Open Access

https://doi.org/10.51173/jt.v5i1.1027

Copy DOI

Journal: Journal of Techniques	Publication Date: Mar 23, 2023
License type: CC BY 4.0

Affiliation: Middle Technical University

Abstract

Recently, researchers intensified their efforts on a dataset with a large number of features named Big Data because of the technological revolution and the development in the data science sector. Dimensionality reduction technology has efficient, effective, and influential methods for analyzing this data, which contains many variables. The importance of Dimensionality Reduction technology lies in several fields, including “data processing, patterns recognition, machine learning, and data mining”. This paper compares two essential methods of dimensionality reduction, Feature Extraction and Feature Selection Which Machine Learning models frequently employ. We applied many classifiers like (Support vector machines, k-nearest neighbors, Decision tree, and Naive Bayes ) to the data of the anthropometric survey of US Army personnel (ANSUR 2) to classify the data and test the relevance of features by predicting a specific feature in USA Army personnel results showing that (k-nearest neighbors) achieved high accuracy (83%) in prediction, then reducing the dimensions by several techniques like (Highly Correlated Filter, Recursive Feature Elimination, and principal components Analysis) results showing that (Recursive Feature Elimination) have the best accuracy by (66%), From these results, it is clear that the efficiency of dimension reduction techniques varies according to the nature of the data. Some techniques are more efficient than others in text data and others are more efficient in dealing with images.

Full Text