Density Measure in Context Clustering for Distributional Semantics of Word Sense Induction

Masood Ghayoomi

doi:10.7508/jist.2020.01.002

Abstract

Word Sense Induction (WSI) aims at inducing word senses from data without using a prior knowledge. Utilizing no labeled data motivated researchers to use clustering techniques for this task. There exist two types of clustering algorithm: parametric or non-parametric. Although non-parametric clustering algorithms are more suitable for inducing word senses, their shortcomings make them useless. Meanwhile, parametric clustering algorithms show competitive results, but they suffer from a major problem that is requiring to set a predefined fixed number of clusters in advance. The main contribution of this paper is to show that utilizing the silhouette score normally used as an internal evaluation metric to measure the clusters’ density in a parametric clustering algorithm, such as K-means, in the WSI task captures words’ senses better than the state-of-the-art models. To this end, word embedding approach is utilized to represent words’ contextual information as vectors. To capture the context in the vectors, we propose two modes of experiments: either using the whole sentence, or limited number of surrounding words in the local context of the target word to build the vectors. The experimental results based on V-measure evaluation metric show that the two modes of our proposed model beat the state-of-the-art models by 4.48% and 5.39% improvement. Moreover, the average number of clusters and the maximum number of clusters in the outputs of our proposed models are relatively equal to the gold data

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Density Measure in Context Clustering for Distributional Semantics of Word Sense Induction

Abstract

Talk to us

Similar Papers

Lead the way for us

Similar Papers

Clustering and Switching Strategies in Verbal Fluency Tasks: Comparison between Amyotrophic Lateral Sclerosis (ALS) and Healthy Controls
...
-
, et. al. ...
01 Mar 2019
01 Mar 2019

Incorporating K-means, Hierarchical Clustering and PCA in Customer Segmentation

-

03 Feb 2021
03 Feb 2021

Extragalactic machine learning : in theory and in practice

-

01 Feb 2021
01 Feb 2021

Word Sense Induction in Persian and English: A Comparative Study
Masood Ghayoomi
Journal of Information Systems and Telecommunication (JIST) | VOL. 9
Masood GhayoomiMasood Ghayoomi
27 Oct 2021
Journal of Information Systems and Telecommunication (JIST) | VOL. 9

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Density Measure in Context Clustering for Distributional Semantics of Word Sense Induction

Abstract

Talk to us

Similar Papers