Abstract

Text clustering has become an important part of the web data organization with the rapid growth of the World Wide Web (www). Clustering simplifies web search engine work by grouping large amount of documents, retrieved according to a given query. Similarity measures used in clustering affect the output of the grouping directly. Most of the document clustering techniques rely on single term analysis of text, such as vector space model. In order to improve grouping of Turkish documents, we investigate several similarity measures based on the semantic similarity of terms. Moreover, some techniques for calculating documents similarity are studied. The aim of this paper is to study the effects of semantic and single term similarity measures to the clustering results of Turkish documents. All experiments are carried out on Turkish web sites, taking into account the relationships of terms based on the ontology for the Turkish language.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.