Named entity recognition in Bengali and Hindi using support vector machine

Asif Ekbal,Sivaji Bandyopadhyay

doi:10.1075/li.34.1.02ekb

Abstract

Named Entity Recognition (NER) aims to classify each word of a document into predefined target named entity (NE) classes and is nowadays considered to be fundamental for many Natural Language Processing (NLP) tasks such as information retrieval, machine translation, information extraction, question answering systems and others. This paper reports about the development of a NER system for Bengali and Hindi using Support Vector Machine (SVM). We have used the annotated corpora of 122,467 tokens of Bengali and 502,974 tokens of Hindi tagged with the twelve different NE classes, defined as part of the IJCNLP-08 NER Shared Task for South and South East Asian Languages (SSEAL). An appropriate tag conversion routine has been developed in order to convert the data into the forms tagged with the four NE tags, namely Person name, Location name, Organization name and Miscellaneous name. The system makes use of the different contextual information of the words along with the variety of orthographic word-level features that are helpful in predicting the different NE classes. The system has been tested with the gold standard test sets of 35K, and 38K tokens for Bengali, and Hindi, respectively. Evaluation results have demonstrated the overall recall, precision, and f-score values of 85.11%, 81.74%, and 83.39%, respectively, for Bengali and 82.76%, 77.81%, and 80.21%, respectively, for Hindi. Statistical analysis, ANOVA is performed to show that the improvement in the performance with the use of language dependent features is statistically significant over the language independent features for Bengali and Hindi both.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Named entity recognition in Bengali and Hindi using support vector machine

Abstract

Talk to us

Similar Papers

More From: Lingvisticæ Investigationes

Lead the way for us

Journal: Lingvisticæ Investigationes	Publication Date: Jul 7, 2011
Citations: 11

Similar Papers

Named Entity Recognition using Support Vector Machine: A Language Independent Approach
...
Zenodo (CERN European Organization for Nuclear Research) | VOL. -
, et. al. ...
23 Mar 2010
Zenodo (CERN European Organization for Nuclear Research) | VOL. -

A Multiengine NER System with Context Pattern Learning and Post-processing Improves System Performance
Asif Ekbal ... Sivaji Bandyopadhyay
International Journal of Computer Processing of Languages | VOL. 22
Asif Ekbal, et. al.Asif Ekbal ... Sivaji Bandyopadhyay
01 Jun 2009
International Journal of Computer Processing of Languages | VOL. 22

Named Entity Recognition in Bengali
Asif Ekbal ... Sivaji Bandyopadhyay
Northern European Journal of Language Technology | VOL. 1
Asif Ekbal, et. al.Asif Ekbal ... Sivaji Bandyopadhyay
02 Feb 2010
Northern European Journal of Language Technology | VOL. 1

A Comparative Study of Dictionary-based and Machine Learning-based Named Entity Recognition in Pashto
Rafiullah Momand ... Shakirullah Waseeb
-
Rafiullah Momand, et. al.Rafiullah Momand ... Shakirullah Waseeb
18 Dec 2020
18 Dec 2020

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Named entity recognition in Bengali and Hindi using support vector machine

Abstract

Talk to us

Similar Papers

More From: Lingvisticæ Investigationes