A Multilingual Datasets Repository of the Hadith Content

Ahsan Mahmood,Fawaz K,Mahwish Ilyas,Hikmat Ullah,Muhammad Ramzan

doi:10.14569/ijacsa.2018.090224

Abstract

Knowledge extraction from unstructured data is a challenging research problem in research domain of Natural Language Processing (NLP). It requires complex NLP tasks like entity extraction and Information Extraction (IE), but one of the most challenging tasks is to extract all the required entities of data in the form of structured format so that data analysis can be applied. Our focus is to explain how the data is extracted in the form of datasets or conventional database so that further text and data analysis can be carried out. This paper presents a framework for Hadith data extraction from the Hadith authentic sources. Hadith is the collection of sayings of Holy Prophet Muhammad, who is the last holy prophet according to Islamic teachings. This paper discusses the preparation of the dataset repository and highlights issues in the relevant research domain. The research problem and their solutions of data extraction, pre-processing and data analysis are elaborated. The results have been evaluated using the standard performance evaluation measures. The dataset is available in multiple languages, multiple formats and is available free of cost for research purposes.

Highlights

Data mining, Information retrieval and knowledge extraction have become attractive fields for the researchers during the last decade due to the birth of Social media [1]
There is a lack of a repository containing data sets of Hadith for researchers to work in various research domains, such as text mining, data analysis, information retrieval and knowledge extraction
It is not possible to propose a single algorithm for data extraction from the Hadith that can be applied to all the Hadith books of different languages

Summary

A Multilingual Datasets Repository of the Hadith Content

Abstract—Knowledge extraction from unstructured data is a challenging research problem in research domain of Natural Language Processing (NLP). It requires complex NLP tasks like entity extraction and Information Extraction (IE), but one of the most challenging tasks is to extract all the required entities of data in the form of structured format so that data analysis can be applied. Our focus is to explain how the data is extracted in the form of datasets or conventional database so that further text and data analysis can be carried out. This paper discusses the preparation of the dataset repository and highlights issues in the relevant research domain. The research problem and their solutions of data extraction, preprocessing and data analysis are elaborated.

INTRODUCTION

BACKGROUND

RESEARCH METHODOLOGY

Selection of Hadith Sources

Hadith Content Extraction

Dataset Preparation

Website Creation

CONCLUSION

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: International Journal of Advanced Computer Science and Applications	Publication Date: Jan 1, 2018
Citations: 12	License type: cc-by

R Discovery Prime

R Discovery Prime

A Multilingual Datasets Repository of the Hadith Content

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications

Lead the way for us

Similar Papers

Examining Knowledge Extraction Processes from Heterogeneous Data Sources
Serdar Kürşat Sarıkoz
Brilliant Engineering | VOL. 4
Serdar Kürşat SarıkozSerdar Kürşat Sarıkoz
08 Feb 2023
Brilliant Engineering | VOL. 4

Natural Language Processing and the Promise of Big Data: Small Step Forward, but Many Miles to Go.
Thomas M Maddox ... Michael A Matheny
Circulation. Cardiovascular quality and outcomes | VOL. 8
Thomas M Maddox, et. al.Thomas M Maddox ... Michael A Matheny
18 Aug 2015
Circulation. Cardiovascular quality and outcomes | VOL. 8

Knowledge extraction from unstructured data and classification through distributed ontologies

-

01 Jan 2012
01 Jan 2012

Named entity recognition and its role in unstructured data analysis
Oleh R Staso ... Nazarii Ye Burak
Informatics. Culture. Technology | VOL. 1
Oleh R Staso, et. al.Oleh R Staso ... Nazarii Ye Burak
26 Sep 2024
Informatics. Culture. Technology | VOL. 1

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Multilingual Datasets Repository of the Hadith Content

Abstract

Highlights

Summary

Talk to us

Similar Papers

More From: International Journal of Advanced Computer Science and Applications