Automated hierarchical classification of scanned documents using convolutional neural network and regular expression

Rifiana Arief,Hustinawaty Hustinawaty,Tubagus Maulana Kusuma,Achmad Benny Mutiara

doi:10.11591/ijece.v12i1.pp1018-1029

Abstract

<p>This research proposed automated hierarchical classification of scanned documents with characteristics content that have unstructured text and special patterns (specific and short strings) using convolutional neural network (CNN) and regular expression method (REM). The research data using digital correspondence documents with format PDF images from pusat data teknologi dan informasi (technology and information data center). The document hierarchy covers type of letter, type of manuscript letter, origin of letter and subject of letter. The research method consists of preprocessing, classification, and storage to database. Preprocessing covers extraction using Tesseract optical character recognition (OCR) and formation of word document vector with Word2Vec. Hierarchical classification uses CNN to classify 5 types of letters and regular expression to classify 4 types of manuscript letter, 15 origins of letter and 25 subjects of letter. The classified documents are stored in the Hive database in Hadoop big data architecture. The amount of data used is 5200 documents, consisting of 4000 for training, 1000 for testing and 200 for classification prediction documents. The trial result of 200 new documents is 188 documents correctly classified and 12 documents incorrectly classified. The accuracy of automated hierarchical classification is 94%. Next, the search of classified scanned documents based on content can be developed.</p>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Journal: International Journal of Electrical and Computer Engineering (IJECE)	Publication Date: Feb 1, 2022
Citations: 3	License type: CC BY-SA 4.0

R Discovery Prime

R Discovery Prime

Automated hierarchical classification of scanned documents using convolutional neural network and regular expression

Abstract

Talk to us

Similar Papers

More From: International Journal of Electrical and Computer Engineering (IJECE)

Lead the way for us

Similar Papers

Research on improved convolutional wavelet neural network
Jingwei Liu ... Jiaxin Li
Scientific Reports | VOL. 11
Jingwei Liu, et. al.Jingwei Liu ... Jiaxin Li
09 Sep 2021
Scientific Reports | VOL. 11

디지털 이미지에 기록된 시간 정보를 추출하기 위한 OCR 기법들에 대한 분석
Yong Jin Kim ... Nam In Park
Korean Journal of Forensic Science | VOL. 22
Yong Jin Kim, et. al.Yong Jin Kim ... Nam In Park
30 Nov 2021
Korean Journal of Forensic Science | VOL. 22

An OCR Engine for Printed Receipt Images using Deep Learning Techniques
Cagri Sayallar ... Nurcan Babalik
International Journal of Advanced Computer Science and Applications | VOL. 14
Cagri Sayallar, et. al.Cagri Sayallar ... Nurcan Babalik
01 Jan 2023
International Journal of Advanced Computer Science and Applications | VOL. 14

Digitization of Data from Invoice using OCR
Venkata Naga Sai Rakesh Kamisetty ... L Mary Gladence
-
Venkata Naga Sai Rakesh Kamisetty, et. al.Venkata Naga Sai Rakesh Kamisetty ... L Mary Gladence
29 Mar 2022
29 Mar 2022

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Automated hierarchical classification of scanned documents using convolutional neural network and regular expression

Abstract

Talk to us

Similar Papers

More From: International Journal of Electrical and Computer Engineering (IJECE)