Simplifying the reading of historical manuscripts

Abedelkadir Asi,Jihad El-Sana,Rafi Cohen,Klara Kedem

doi:10.1109/icdar.2015.7333877

Abstract

Complex document layouts pose prominent challenges for document image understanding algorithms. These layouts impose irregularities on the location of text paragraphs which consequently induces difficulties in reading the text. In this paper we present a robust framework for analyzing historical manuscripts with complex layouts. This framework aims to provide a convenient reading experience for historians through topnotch algorithms for text localization, classification and dewarping. We segment text into spatially coherent regions and text-lines using texture-based filters and refine this segmentation by exploiting Markov Random Fields (MRFs). A principled technique is presented for dewarping curvy text regions using a non-linear geometric transformation. The framework has been validated using a subset of a publicly available dataset of historical documents and it provided promising results.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Simplifying the reading of historical manuscripts

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Designing Explainable Text Classification Pipelines: Insights from IT Ticket Complexity Prediction Case Study
Aleksandra Revina ... Krisztian Buza
-
Aleksandra Revina, et. al.Aleksandra Revina ... Krisztian Buza
01 Jan 2020
01 Jan 2020

Color segmentation for historical documents using Markov random fields
Werner Pantke ... Arne Haak
-
Werner Pantke, et. al.Werner Pantke ... Arne Haak
01 Aug 2014
01 Aug 2014

A Combined Algorithm for Layout Analysis of Arabic Document Images and Text Lines Extraction
Abdulrahman Alshameri ... Sherif Abdou
International Journal of Computer Applications | VOL. 49
Abdulrahman Alshameri, et. al.Abdulrahman Alshameri ... Sherif Abdou
31 Jul 2012
International Journal of Computer Applications | VOL. 49

A Multi-task Text Classification Model Based on Label Embedding Learning
Yuemei Xu ... Zuwei Fan
-
Yuemei Xu, et. al.Yuemei Xu ... Zuwei Fan
01 Jan 2021
01 Jan 2021

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Simplifying the reading of historical manuscripts

Abstract

Talk to us

Similar Papers