Generating Javanese Stopwords List using K-means Clustering Algorithm

Aji Prasetya Wibawa,Hidayah Kariima Fithri,Andrew Nafalski,Ilham Ari Elbaith Zaeni

doi:10.17977/um018v3i22020p106-111

Generating Javanese Stopwords List using K-means Clustering Algorithm

Aji Prasetya Wibawa, Hidayah Kariima Fithri + Show 2 more

Open Access

https://doi.org/10.17977/um018v3i22020p106-111

Copy DOI

Journal: Knowledge Engineering and Data Science	Publication Date: Dec 31, 2020
Citations: 4	License type: CC BY-SA 4.0

Affiliation: State University of Malang, University of South Australia

#Stopword List #K-means Clustering Method + Show 8 more

Abstract
Full-Text PDF
Similar Papers

Abstract

Stopword removal necessary in Information Retrieval. It can remove frequently appeared and general words to reduce memory storage. The algorithm eliminates each word that is precisely the same as the word in the stopword list. However, generating the list could be time-consuming. The words in a specific language and domain must be collected and validated by specialists. This research aims to develop a new way to generate a stop word list using the K-means Clustering method. The proposed approach groups words based on their frequency. The confusion matrix calculates the difference between the findings with a valid stopword list created by a Javanese linguist. The accuracy of the proposed method is 78.28% (K=7). The result shows that the generation of Javanese stopword lists using a clustering method is reliable.

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Similar Papers

Paper Title

Journal

Date

Author

View more papers

More From: Knowledge Engineering and Data Science

Paper Title

Journal

Date

Author

View more papers

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.