Abstract

Document Layout Analysis (DLA) is a segmentation process that decomposes a scanned document image into its blocks of interest and classifies them. DLA is essential in a large number of applications, such as Information Retrieval, Machine Translation, Optical Character Recognition (OCR) systems, and structured data extraction from documents. However, identification of document blocks in DLA is challenging due to variations of block locations, inter-and intra-class variability, and background noises. In this paper, we propose a novel texture-based convolutional neural network for document layout analysis, called DoT-Net. DoT-Net is a multiclass classifier that can effectively identify document component blocks such as text, image, table, mathematical expression, and line-diagram, whereas most related methods have focused on the text vs. non-text block classification problem. DoT-Net can capture textural variations among the multiclass regions of documents. Our proposed method DoT-Net achieved promising results outperforming state-of-the-art document layout classifiers on accuracy, F1 score, and AUC. The open-source code of DoT-Net is available at https://github.com/datax-lab/DoTNet.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.