Performance Assessment of Multiple Classifiers Based on Ensemble Feature Selection Scheme for Sentiment Analysis

Monalisa Ghosh,Goutam Sanyal

doi:10.1155/2018/8909357

Abstract

Sentiment classification or sentiment analysis has been acknowledged as an open research domain. In recent years, an enormous research work is being performed in these fields by applying various numbers of methodologies. Feature generation and selection are consequent for text mining as the high-dimensional feature set can affect the performance of sentiment analysis. This paper investigates the inability or incompetency of the widely used feature selection methods (IG, Chi-square, and Gini Index) with unigram and bigram feature set on four machine learning classification algorithms (MNB, SVM, KNN, and ME). The proposed methods are evaluated on the basis of three standard datasets, namely, IMDb movie review and electronics and kitchen product review dataset. Initially, unigram and bigram features are extracted by applying n-gram method. In addition, we generate a composite features vector CompUniBi (unigram + bigram), which is sent to the feature selection methods Information Gain (IG), Gini Index (GI), and Chi-square (CHI) to get an optimal feature subset by assigning a score to each of the features. These methods offer a ranking to the features depending on their score; thus a prominent feature vector (CompIG, CompGI, and CompCHI) can be generated easily for classification. Finally, the machine learning classifiers SVM, MNB, KNN, and ME used prominent feature vector for classifying the review document into either positive or negative. The performance of the algorithm is measured by evaluation methods such as precision, recall, and F-measure. Experimental results show that the composite feature vector achieved a better performance than unigram feature, which is encouraging as well as comparable to the related research. The best results were obtained from the combination of Information Gain with SVM in terms of highest accuracy.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: Applied Computational Intelligence and Soft Computing	Publication Date: Oct 1, 2018
Citations: 17	License type: CC BY 4.0

R Discovery Prime

R Discovery Prime

Performance Assessment of Multiple Classifiers Based on Ensemble Feature Selection Scheme for Sentiment Analysis

Abstract

Talk to us

Similar Papers

More From: Applied Computational Intelligence and Soft Computing

Lead the way for us

Similar Papers

Analysing sentiments based on multi feature combination with supervised learning
Monalisha Ghosh ... Goutam Sanyal
International Journal of Data Mining, Modelling and Management | VOL. 11
Monalisha Ghosh, et. al.Monalisha Ghosh ... Goutam Sanyal
01 Jan 2019
International Journal of Data Mining, Modelling and Management | VOL. 11

Analysing sentiments based on multi feature combination with supervised learning
Monalisha Ghosh ... Goutam Sanyal
International Journal of Data Mining, Modelling and Management | VOL. 11
Monalisha Ghosh, et. al.Monalisha Ghosh ... Goutam Sanyal
01 Jan 2019
International Journal of Data Mining, Modelling and Management | VOL. 11

Multi-class SVM Classification Comparison for Health Service Satisfaction Survey Data in Bahasa
Gede Indrawan ... Aris Gunadi
HighTech and Innovation Journal | VOL. 3
Gede Indrawan, et. al.Gede Indrawan ... Aris Gunadi
01 Dec 2022
HighTech and Innovation Journal | VOL. 3

An ensemble approach to stabilize the features for multi-domain sentiment analysis using supervised machine learning
Monalisa Ghosh ... Goutam Sanyal
Journal of Big Data | VOL. 5
Monalisa Ghosh, et. al.Monalisa Ghosh ... Goutam Sanyal
14 Nov 2018
Journal of Big Data | VOL. 5

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Performance Assessment of Multiple Classifiers Based on Ensemble Feature Selection Scheme for Sentiment Analysis

Abstract

Talk to us

Similar Papers

More From: Applied Computational Intelligence and Soft Computing