Extraction of text content from PDF documents based on automaton theory

Xiao-Juan Wang,Jin-Gang Liu,Yan-Bing Liu,Jian-Long Tan

doi:10.3724/sp.j.1087.2012.02491

Extraction of text content from PDF documents based on automaton theory

Xiao-Juan Wang, Jin-Gang Liu + Show 2 more

https://doi.org/10.3724/sp.j.1087.2012.02491

Copy DOI

Journal: Journal of Computer Applications

Publication Date: May 13, 2013

#Deterministic Finite Automaton #PDF File + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

The existing methods of extracting text content from a PDF file,such as the one adopted by the PDFBox library,are not efficient enough to handle the high-speed network traffic.Moreover,these methods cannot extract the contents streamingly from partial PDF packets in transfer.This paper proposed a new method based on automaton theory.The method adopted a hierarchical keyword Deterministic Finite Automaton(DFA) to extract information from complete or incomplete PDF files.The experimental results show that the response time of the proposed method is about 17%-37% of the algorithm used by PDFBox when processing PDF files in Chinese or English.

Full Text

Paper version not known

Open DOI Link

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: Journal of Computer Applications

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.