A Benchmark Kannada Handwritten Document Dataset and Its Segmentation

Alireza Alaei,Umapada Pal,P Nagabhushan

doi:10.1109/icdar.2011.37

Abstract

Research towards Indian handwritten document analysis achieved increasing attention in recent years. In pattern recognition and especially in handwritten document recognition, standard databases play vital roles for evaluating performances of algorithms and comparing results obtained by different groups of researchers. For Indian languages, there is a lack of standard database of handwritten texts to evaluate performance of different document recognition approaches and for comparison purpose. In this paper, an unconstrained Kannada handwritten text database (KHTD) is introduced. The KHTD contains 204 handwritten documents of four different categories written by 51 native speakers of Kannada. Total number of text-lines and words in the dataset are 4298 and 26115, respectively. In most of text-pages of the KHTD contains either an overlapping or a touching text-lines and the average number of text-lines in each document on the database is 21. Two types of ground truths based on pixels information and content information are generated for the database. Providing these two types of ground truths for the KHTD, it can be utilized in many areas of document image processing such as sentence recognition/understanding, text-line segmentation, word segmentation, word recognition, and character segmentation. To provide a framework for other researches, recent text-line segmentation results on this dataset are also reported. The KHTD is available for research purposes.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

A Benchmark Kannada Handwritten Document Dataset and Its Segmentation

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

A New Dataset of Persian Handwritten Documents and Its Segmentation
Alireza Alaei ... P Nagabhushan
-
Alireza Alaei, et. al.Alireza Alaei ... P Nagabhushan
01 Nov 2011
01 Nov 2011

DATASET AND GROUND TRUTH FOR HANDWRITTEN TEXT IN FOUR DIFFERENT SCRIPTS
Alireza Alaei ... Umapada Pal
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 26
Alireza Alaei, et. al.Alireza Alaei ... Umapada Pal
01 Jun 2012
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 26

A hierarchical approach to recognition of handwritten Bangla characters
Subhadip Basu ... Dipak Kumar Basu
Pattern Recognition | VOL. 42
Subhadip Basu, et. al.Subhadip Basu ... Dipak Kumar Basu
17 Jan 2009
Pattern Recognition | VOL. 42

A robust method for line and word segmentation in handwritten text
Abdelaali Hassaine
-
Abdelaali HassaineAbdelaali Hassaine
01 Jan 2013
01 Jan 2013

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

A Benchmark Kannada Handwritten Document Dataset and Its Segmentation

Abstract

Talk to us

Similar Papers