Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

Aysha Akther,Kazi Masudul Alam,Rameswar Debnath,Hafsa Sultana,Sujana Saha,A K Z Rasel Rahman,Md Shymon Islam

doi:10.1109/access.2022.3195236

Aysha Akther, Kazi Masudul Alam + Show 5 more

Open Access

https://doi.org/10.1109/access.2022.3195236

Copy DOI

Journal: IEEE Access	Publication Date: Jan 1, 2022
Citations: 6	License type: CC BY-NC-ND 4.0

Affiliation: Khulna University

Abstract

Research in Natural Language Processing (NLP) and computational linguistics highly depends on a good quality representative corpus of any specific language. Bangla is one of the most spoken languages in the world but Bangla NLP research is in its early stage of development due to the lack of quality public corpus. This article describes the detailed compilation methodology of a comprehensive monolingual Bangla corpus, KUMono. The newly developed corpus consists of more than 350 million word tokens and more than one million unique tokens from 18 major text categories of online Bangla websites. We have conducted several word-level and character-level linguistic phenomenon analyses based on empirical studies of the developed corpus. The corpus follows Zipf’s curve and hapax legomena rule. The quality of the corpus is also assessed by analyzing and comparing the inherent sparseness of the corpus with existing Bangla corpora, by analyzing the distribution of function words of the corpus and vocabulary growth rate. We have developed a Bangla article categorization application based on the KUMono corpus and received compelling results by comparing to the state-of-the-art models.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

Abstract

Talk to us

Similar Papers

More From: IEEE Access

Lead the way for us

Similar Papers

Special issue on statistical learning of natural language structured input and output
Lluís Màrquez ... Alessandro Moschitti
Natural Language Engineering | VOL. 18
Lluís Màrquez, et. al.Lluís Màrquez ... Alessandro Moschitti
14 Mar 2012
Natural Language Engineering | VOL. 18

From semantics to pragmatics: where IS can lead in Natural Language Processing (NLP) research
Yan Li ... Dapeng Liu
European Journal of Information Systems | VOL. 30
Yan Li, et. al.Yan Li ... Dapeng Liu
24 Sep 2020
European Journal of Information Systems | VOL. 30

A two-site survey of medical center personnel\u2019s willingness to share clinical data for research: implications for reproducible health NLP research
Chunhua Weng ... John F Hurdle
BMC Medical Informatics and Decision Making | VOL. 19
Chunhua Weng, et. al.Chunhua Weng ... John F Hurdle
01 Apr 2019
BMC Medical Informatics and Decision Making | VOL. 19

Natural Language Processing and Computational Linguistics
Junichi Tsujii
Computational Linguistics | VOL. -
Junichi TsujiiJunichi Tsujii
07 Dec 2021
Computational Linguistics | VOL. -

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Compilation, Analysis and Application of a Comprehensive Bangla Corpus KUMono

Abstract

Talk to us

Similar Papers

More From: IEEE Access