Abstract
The feature selection is an important part in automatic text classification. In this paper, we use a Chinese semantic dictionary -- Hownet to extract the concepts from the word as the feature set, because it can better reflect the meaning of the text. We construct a combined feature set that consists of both sememes and the Chinese words, propose a CHI-MCOR weighing method according to the weighing theories and classification precision. The effectiveness of the competitive network and the Radial Basis Function (RBF) network in text classification are examined. Experimental result shows that if the words are extracted properly, not only the feature dimension is smaller but also the classification precision is higher, the RBF network outperform competitive network for automatic text classification because of the application of supervised learning. Besides its much shorter training time than the BP network's, the RBF network makes precision and recall rates that are almost at the same level as the BP network's.
Talk to us
Join us for a 30 min session where you can share your feedback and ask us any queries you have
Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.