Usable Amharic text corpus for natural language processing applications

Michael Melese Woldeyohannis,Million Meshesha

doi:10.1016/j.acorp.2022.100033

Abstract

In this paper, we describe the preparation of a usable Amharic text corpus for different Natural Language Processing (NLP) applications. Natural language applications, such as document classification, topic modeling, machine translation, speech recognition, and others, suffer greatly from a lack of digital resources. This is especially true for Amharic, a resource-constrained, morphologically rich, and complex language. In response to this, a total of 67,739 Amharic news documents consisting of 8 different categories from online sources are collected. The collected corpus passes through a number of pre-processing steps including; data cleaning, text normalization and punctuation correction. To validate the usability of the collected corpora from different domains, a baseline document classification experiment was conducted. Experimental results show that, 84.53% accuracy is registered using deep learning in the absence of linguistic information. Finding indicated that it is possible to use the prepared corpora for different natural language applications in the absence of linguistic resources such as stemmer and dictionary despite the complexity of Amharic language. We are further working towards Amharic news document classification by incorporating a linguistic independent stop-word detection, stemming and unsupervised morphological segmentation of Amharic documents.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Usable Amharic text corpus for natural language processing applications

Abstract

Talk to us

Similar Papers

More From: Applied Corpus Linguistics

Lead the way for us

Journal: Applied Corpus Linguistics	Publication Date: Dec 1, 2022
Citations: 1

Similar Papers

The Application of Natural Language Processing and Automated Scoring in Second Language Assessment

-

22 Dec 2012
22 Dec 2012

Guest Editors Introduction: Machine Learning in Speech and Language Technologies
Pascale Fung ... Dan Roth
Machine Learning | VOL. 60
Pascale Fung, et. al.Pascale Fung ... Dan Roth
01 Sep 2005
Machine Learning | VOL. 60

Machine Learning for Text
Charu C Aggarwal
-
Charu C AggarwalCharu C Aggarwal
01 Jan 2018
01 Jan 2018

Neural Natural Language Generation: A Survey on Multilinguality, Multimodality, Controllability and Learning
Erkut Erdem ... Barbara Plank
Journal of Artificial Intelligence Research | VOL. 73
Erkut Erdem, et. al.Erkut Erdem ... Barbara Plank
06 Apr 2022
Journal of Artificial Intelligence Research | VOL. 73

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Usable Amharic text corpus for natural language processing applications

Abstract

Talk to us

Similar Papers

More From: Applied Corpus Linguistics