Unsupervised deep learning for text line segmentation

Berat Kurar

doi:10.48448/zf1x-zh06

Abstract

We present an unsupervised deep learning method for text line segmentation that is inspired by the relative variance between text lines and spaces among text lines. Handwritten text line segmentation is important for the efficiency of further processing. A common method is to train a deep learning network for embedding the document image into an image of blob lines that are tracing the text lines. Previous methods learned such embedding in a supervised manner, requiring the annotation of many document images. This paper presents an unsupervised embedding of document image patches without a need for annotations. The number of foreground pixels over the text lines is relatively different from the number of foreground pixels over the spaces among text lines. Generating similar and different pairs relying on this principle definitely leads to outliers. However, as the results show, the outliers do not harm the convergence and the network learns to discriminate the text lines from the spaces between text lines. Remarkably, with a challenging Arabic handwritten text line segmentation dataset, VML-AHTE, we achieved superior performance over the supervised methods. Additionally, the proposed method was evaluated on the ICDAR 2017 and ICFHR 2010 handwritten text line segmentation datasets.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Unsupervised deep learning for text line segmentation

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Unsupervised deep learning for text line segmentation
Berat Kurar Barakat ... Reem Alaasam
-
Berat Kurar Barakat, et. al.Berat Kurar Barakat ... Reem Alaasam
10 Jan 2021
10 Jan 2021

Unsupervised Learning of Text Line Segmentation by Differentiating Coarse Patterns
Berat Kurar Barakat ... Jihad El-Sana
-
Berat Kurar Barakat, et. al.Berat Kurar Barakat ... Jihad El-Sana
01 Jan 2020
01 Jan 2020

A Novel Text Line Segmentation Method Based on Contour Curve Tracking for Tibetan Historical Documents
Fengming Zhou ... Qiang Lin
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 32
Fengming Zhou, et. al.Fengming Zhou ... Qiang Lin
20 Jun 2018
International Journal of Pattern Recognition and Artificial Intelligence | VOL. 32

A novel method of text line segmentation for historical document image of the uchen Tibetan
Zhenjiang Li ... Yusheng Hao
Journal of Visual Communication and Image Representation | VOL. 61
Zhenjiang Li, et. al.Zhenjiang Li ... Yusheng Hao
01 Mar 2019
Journal of Visual Communication and Image Representation | VOL. 61

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Unsupervised deep learning for text line segmentation

Abstract

Talk to us

Similar Papers