Building a Sentiment Analysis System Using Automatically Generated Training Dataset

Daoud M Daoud,Samir Abou El-Seoud

doi:10.3991/ijoe.v16i06.13623

Abstract

In this paper, we describe a methodology to develop a large training set for sentiment analysis automatically. We extract Arabic tweets and then annotates them for negativeness and positiveness sentiment without human intervention. These annotated tweets are used as a training data set to build our experimental sentiment analysis by using Naive Bayes algorithm and TF-IDF enhancement. The large size of training data for a highly inflected language is necessary to compensate for the sparseness nature of such languages. We present our techniques and explain our experimental system. We use 200 thousand annotated tweets to train our system. The evaluation shows that our sentiment analysis system has high precision and accuracy measures compared to existing ones.

Highlights

Opinions are difficult to extract and search using the current information retrieval relevancy methods
Concerning Arabic sentiment analysis, machine learning techniques are used more than other approaches such as Lexicon based and hybrid techniques
Other researchers have recently processing text collected from social media and designed to handle both Arabic dialects as well as MSA text

Summary

Introduction

Opinions are difficult to extract and search using the current information retrieval relevancy methods. It is even more challenging for languages with little resources such as Arabic. Sentiment analysis is becoming an important mean which helps by automating the process of extracting the opinions from diverse content channels. To handle the richness of Arabic and its inflection nature, we need a training set consisting of thousands of tweets that are rich with negative and positive terms. Such a training dataset does not exist [3, 4]. We present the results of the evaluation of our sentiment analysis system

Related Work

Proposed Model

Tokenization

Stemming

Transformation and building the vector space model

Naive Bayes training algorithm

Preparing the Training Dataset

Conducting the Experiment

Evaluation and Results

Authors

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Building a Sentiment Analysis System Using Automatically Generated Training Dataset

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Online and Biomedical Engineering (iJOE)

Lead the way for us

Journal: International Journal of Online and Biomedical Engineering (iJOE)	Publication Date: May 28, 2020
License type: cc-by

Similar Papers

Building a Sentiment Analysis system using automatically generated training Dataset
Daoud M Daoud ... M Samir Abou Oudi
-
Daoud M Daoud, et. al.Daoud M Daoud ... M Samir Abou Oudi
09 Apr 2019
09 Apr 2019

BiasFinder: Metamorphic Test Generation to Uncover Bias for Sentiment Analysis Systems
Muhammad Hilmi Asyrofi ... David Lo
IEEE Transactions on Software Engineering | VOL. -
Muhammad Hilmi Asyrofi, et. al.Muhammad Hilmi Asyrofi ... David Lo
01 Jan 2020
IEEE Transactions on Software Engineering | VOL. -

BiasHeal: On-the-Fly Black-Box Healing of Bias in Sentiment Analysis Systems
Zhou Yang ... Harshit Jain
-
Zhou Yang, et. al.Zhou Yang ... Harshit Jain
01 Sep 2021
01 Sep 2021

A novel classification approach based on Naïve Bayes for Twitter sentiment analysis
...
KSII Transactions on Internet and Information Systems | VOL. 11
, et. al. ...
30 Jun 2017
KSII Transactions on Internet and Information Systems | VOL. 11

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Building a Sentiment Analysis System Using Automatically Generated Training Dataset

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Online and Biomedical Engineering (iJOE)