Compression-based arabic text classification

Haneen Ta'Amneh,Ehsan Abu Keshek,Yaser Jararweh,Manar Bani Issa,Mahmoud Al-Ayyoub

doi:10.1109/aiccsa.2014.7073253

Compression-based arabic text classification

Haneen Ta'Amneh, Ehsan Abu Keshek + Show 3 more

https://doi.org/10.1109/aiccsa.2014.7073253

Copy DOI

Publication Date: Nov 1, 2014

Citations: 15

Affiliation: Jordan University of Science and Technology

#Text Classification #Problems In Text Mining + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

Text classification (TC) is one of the fundamental problems in text mining. Plenty of works exist on TC with interesting approaches and excellent results; however, most of these works follow a word-based approach for feature extraction. In this work, we are interested in an alternative (byte-based or character-based) approach known as compression-based TC (CTC). CTC has been used for some languages such as English and Portuguese and it is shown to have certain advantages/ disadvantages compared with word-based approaches. This work applies CTC on the Arabic language with the purpose of investigating whether these advantages/disadvantages exists for the Arabic language as well. The results are encouraging as they show the viability of using CTC for Arabic TC.

Full Text