Compression and String Matching Method for Printed Document Images

Hajime Imura,Yuzuru Tanaka

doi:10.1109/icdar.2009.182

Abstract

This paper describes a compression technique for printed document images and string matching method on the compressed images.To send digitized document images over the Web, compression of the document images is required. Moreover, in order to deal with historical letterpress printing collections, it is important to provide a full-text search method for them.The proposed compression scheme is based on character pattern matching \& substitution approach using a string matching technique of document images.The proposed string matching method is independent from the difference of languages and fonts because it uses the pseudo-coding that is based on statistical character shape features.We also use the pseudo-codes in a string matching of compressed documents.The system is as fast as the full-text search of machine-readable texts.Our method was evaluated in the compressed size, calculating recall-precision curves for n-gram-based query strings.The experiments have shown that about 100 pages of document in gray-scale at 300 dpi can be compressed down to around one megabyte.

Full Text