Feature extraction and selection for Arabic tweets authorship authentication

Mahmoud Al-Ayyoub,Yaser Jararweh,Abdullateef Rabab’Ah,Monther Aldwairi

doi:10.1007/s12652-017-0452-1

Abstract

In tweet authentication, we are concerned with correctly attributing a tweet to its true author based on its textual content. The more general problem of authenticating long documents has been studied before and the most common approach relies on the intuitive idea that each author has a unique style that can be captured using stylometric features (SF). Inspired by the success of modern automatic document classification problem, some researchers followed the Bag-Of-Words (BOW) approach for authenticating long documents. In this work, we consider both approaches and their application on authenticating tweets, which represent additional challenges due to the limitation in their sizes. We focus on the Arabic language due to its importance and the scarcity of works related on it. We create different sets of features from both approaches and compare the performance of different classifiers using them. We experiment with various feature selection techniques in order to extract the most discriminating features. To the best of our knowledge, this is the first study of its kind to combine these different sets of features for authorship analysis of Arabic tweets. The results show that combining all the feature sets we compute yields the best results.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Feature extraction and selection for Arabic tweets authorship authentication

Abstract

Talk to us

Similar Papers

More From: Journal of Ambient Intelligence and Humanized Computing

Lead the way for us

Journal: Journal of Ambient Intelligence and Humanized Computing	Publication Date: Feb 7, 2017
Citations: 28

Similar Papers

Authorship attribution of Arabic tweets
Abdullateef Rabab'Ah ... Mahmoud Al-Ayyoub
-
Abdullateef Rabab'Ah, et. al.Abdullateef Rabab'Ah ... Mahmoud Al-Ayyoub
01 Nov 2016
01 Nov 2016

AuthCom: Authorship verification and compromised account detection in online social networks using AHP-TOPSIS embedded profiling based technique
Ravneet Kaur ... Harish Kumar
Expert Systems with Applications | VOL. 113
Ravneet Kaur, et. al.Ravneet Kaur ... Harish Kumar
05 Jul 2018
Expert Systems with Applications | VOL. 113

Effects of Light Stemming on Feature Extraction and Selection for Arabic Documents Classification
Yousif A Alhaj ... Abdelghani Dahou
-
Yousif A Alhaj, et. al.Yousif A Alhaj ... Abdelghani Dahou
30 Nov 2019
30 Nov 2019

Outlier detection using flexible categorization and interrogative agendas
Marcel Boersma ... Nachoem Wijnberg
Decision Support Systems | VOL. 180
Marcel Boersma, et. al.Marcel Boersma ... Nachoem Wijnberg
19 Feb 2024
Decision Support Systems | VOL. 180

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Feature extraction and selection for Arabic tweets authorship authentication

Abstract

Talk to us

Similar Papers

More From: Journal of Ambient Intelligence and Humanized Computing