Phrase processing methods for Japanese text retrieval

Noriko Kando,Masaharu Yoshioka,Kyo Kageura,Keizo Oyama

doi:10.1145/305110.305120

Abstract

This paper examines the effectiveness of different phrase identification and weighting methods for Japanese text retrieval in an operational information retrieval (IR) system, called NACSIS-IR . Based on our previous experiments, we used character-based indexing with positional information and word-or phrase-based query processing, which allowed us to implement sophisticated linguistic analysis on large-scale databases while maintaining adequate efficiency. The results of retrieval experiments on a large-scale Japanese test collection showed that the combination of enhanced phrase identification using patterns defined over part-of-speech tags and our algorithms Phrase2 and Phrase5 made a significant positive contribution to retrieval effectiveness. The paper also discusses indexing and phrase processing of Japanese or East Asian languages.

Full Text