Writing type, script and language identification in heterogeneous documents

Anis Mezghani,Monji Kherallah,Fouad Slimane

doi:10.1504/ijista.2017.10006001

Abstract

In this paper, we propose a writing type, script and language text classification method to automatically determine the identity of texts segmented from heterogeneous document images. These documents are written in Arabic, French and English languages with mixed machine-printed and handwritten text. To handle such a problem, we treat each text-line/word image with a fixed-length sliding window. Each window is represented with 23 simple and efficient features to achieve the writing type and the script identification goal using Gaussian mixture models (GMM). The proposed approach for language identification is based on a bi-gram analysis of an optical character recognition (OCR) output. Experiments have been conducted with handwritten and machine-printed text-blocks, text-lines and words extracted from the Maurdor database. The results reveal the feasibility of our proposed method in writing type, script and language identification.

Full Text