Semantic Based Text Similarity Computation

Yaqi Liu,Zhijiang Li

doi:10.1007/978-981-10-3530-2_43

Yaqi Liu, Zhijiang Li

https://doi.org/10.1007/978-981-10-3530-2_43

Copy DOI

Export

Save

Cite

Publication Date: Jan 1, 2017

Citations: 2

Affiliation: Wuhan University

Abstract
Full-Text
Similar Papers

Abstract

Listen

Text similarity algorithm is widely used in plurality fields, such as copy detection, text classification, machine translation, intelligent question answering system and natural language processing. At present, vector space model algorithm, which is more commonly used, does not consider the information of semantic features adequately, and the accuracy of the semantic similarity computation results can be further improved. This paper proposes a text similarity computation method which combines the HowNet with vector space model. Similarity computation is divided into two levels. In the level of words, words-similarity calculation based on HowNet prevents the loss of semantic information. In the level of texts, text-similarity calculation by vector space model ensures the integrity of the information expressed in the texts. This paper designs an experiment of news text classification based on KNN algorithm, in which data obtained from a part of the Chinese news in Sogou data corpora. Experimental results show that the method proposed in this paper is more accurate than the traditional vector space model algorithm.

Full Text