Multioriented and Curved Text Lines Extraction From Indian Documents

U Pal,P.P Roy

doi:10.1109/tsmcb.2004.827613

Abstract

There are printed artistic documents where text lines of a single page may not be parallel to each other. These text lines may have different orientations or the text lines may be curved shapes. For the optical character recognition (OCR) of these documents, we need to extract such lines properly. In this paper, we propose a novel scheme, mainly based on the concept of water reservoir analogy, to extract individual text lines from printed Indian documents containing multioriented and/or curve text lines. A reservoir is a metaphor to illustrate the cavity region of a character where water can be stored. In the proposed scheme, at first, connected components are labeled and identified either as isolated or touching. Next, each touching component is classified either straight type (S-type) or curve type (C-type), depending on the reservoir base-area and envelope points of the component. Based on the type (S-type or C-type) of a component two candidate points are computed from each touching component. Finally, candidate regions (neighborhoods of the candidate points) of the candidate points of each component are detected and after analyzing these candidate regions, components are grouped to get individual text lines.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Multioriented and Curved Text Lines Extraction From Indian Documents

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Systems, Man and Cybernetics, Part B (Cybernetics)

Lead the way for us

Journal: IEEE Transactions on Systems, Man and Cybernetics, Part B (Cybernetics)	Publication Date: Aug 1, 2004
Citations: 84

Similar Papers

Multi-Oriented English Text Line Extraction Using Background and Foreground Information
Partha Pratim Roy ... Josep Lladós
-
Partha Pratim Roy, et. al.Partha Pratim Roy ... Josep Lladós
01 Sep 2008
01 Sep 2008

Text line extraction in graphical documents using background and foreground information
Partha Pratim Roy ... Umapada Pal
International Journal on Document Analysis and Recognition (IJDAR) | VOL. 15
Partha Pratim Roy, et. al.Partha Pratim Roy ... Umapada Pal
30 Jun 2011
International Journal on Document Analysis and Recognition (IJDAR) | VOL. 15

Text Line Identification from a Multilingual Document
P.A Vijaya ... M.C Padma
-
P.A Vijaya, et. al.P.A Vijaya ... M.C Padma
01 Mar 2009
01 Mar 2009

Multi-oriented Text Recognition in Graphical Documents Using HMM
Partha Pratim Roy ... Umapada Pal
-
Partha Pratim Roy, et. al.Partha Pratim Roy ... Umapada Pal
01 Apr 2014
01 Apr 2014

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Multioriented and Curved Text Lines Extraction From Indian Documents

Abstract

Talk to us

Similar Papers

More From: IEEE Transactions on Systems, Man and Cybernetics, Part B (Cybernetics)