Short text classification based on Wikipedia and Word2vec

Liu Wensen Liu Wensen,Wang Jun Wang Jun,Wang Xiaoyi Wang Xiaoyi,Cao Zewen Cao Zewen

doi:10.1109/compcomm.2016.7924894

Liu Wensen Liu Wensen, Wang Jun Wang Jun + Show 2 more

https://doi.org/10.1109/compcomm.2016.7924894

Copy DOI

Abstract

Different from long texts, the features of Chinese short texts is much sparse, which is the primary cause of the low accuracy in the classification of short texts by using traditional classification methods. In this paper, a novel method was proposed to tackle the problem by expanding the features of short text based on Wikipedia and Word2vec. Firstly, build the semantic relevant concept sets of Wikipedia. We get the articles that have high relevancy with Wikipedia concepts and use the word2vec tools to measure the semantic relatedness between target concepts and related concepts. And then we use the relevant concept sets to extend the short texts. Compared to traditional similarity measurement between concepts using statistical method, this method can get more accurate semantic relatedness. The experimental results show that by expanding the features of short texts, the classification accuracy can be improved. Specifically, our method appeared to be more effective.

Full Text