A Similarity Rough Set Model for Document Representation and Document Clustering

Nguyen Chi Thanh,Koichi Yamada,Muneyuki Unehara

doi:10.20965/jaciii.2011.p0125

A Similarity Rough Set Model for Document Representation and Document Clustering

Nguyen Chi Thanh, Koichi Yamada + Show 1 more

Open Access

https://doi.org/10.20965/jaciii.2011.p0125

Copy DOI

Journal: Journal of Advanced Computational Intelligence and Intelligent Informatics	Publication Date: Mar 20, 2011
Citations: 6	License type: cc-by-nd

#Tolerance Rough Set Model #Document Clustering + Show 8 more

Abstract
Full-Text
Similar Papers

Abstract

Document clustering is a textmining technique for unsupervised document organization. It helps the users browse and navigate large sets of documents. Ho et al. proposed a Tolerance Rough Set Model (TRSM) [1] for improving the vector space model that represents documents by vectors of terms and applied it to document clustering. In this paper we analyze their model to propose a new model for efficient clustering of documents. We introduce Similarity Rough Set Model (SRSM) as another model for presenting documents in document clustering. The model is evaluated by experiments on test collections. The experiment results show that the SRSM document clusteringmethod outperforms the one with TRSM and the results of SRSM are less affected by the value of parameter than TRSM.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

More From: Journal of Advanced Computational Intelligence and Intelligent Informatics

Paper Title

Journal

Date

Author

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.