Reclust: an efficient clustering algorithm for mixed data based on reclustering and cluster validation

Amala Jayanthi Maria Soosai Arockiam,Elizabeth Shanthi Irudhayaraj

doi:10.11591/ijeecs.v29.i1.pp545-552

Amala Jayanthi Maria Soosai Arockiam, Elizabeth Shanthi Irudhayaraj

Open Access

https://doi.org/10.11591/ijeecs.v29.i1.pp545-552

Copy DOI

Abstract

<span>Clustering is a significant approach in data mining, which seeks to find groups or clusters of data. Both numeric and categorical features are frequently used to define the data in real-world applications. Several different clustering algorithms are proposed for the numerical and categorical datasets. In clustering algorithms, the quality of clustering results is evaluated using cluster validation. This paper proposes an efficient clustering algorithm for mixed numerical and categorical data using re-clustering and cluster validation. Initially, the mixed dataset is clustered with four traditional clustering algorithms like expectation-maximization (EM), hierarchical cluster (HC), k-means (KM), and self-organizing map (SOM). These four algorithms are validated, and the best algorithm is selected for re-clustering. It is an iterative process for improving the quality of cluster results. The incorrectly clustered data is iteratively re-clustered and evaluated based on the cluster validation. The performance of the proposed clustering method is evaluated with a real-time dataset in terms of purity, normalized mutual information, rand index, precision, and recall. The experimental results have shown that the proposed reclust algorithm achieves better performance compared to other clustering algorithms.</span>

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

R Discovery Prime

R Discovery Prime

Reclust: an efficient clustering algorithm for mixed data based on reclustering and cluster validation

Abstract

Talk to us

Similar Papers

More From: Indonesian Journal of Electrical Engineering and Computer Science

Lead the way for us

Journal: Indonesian Journal of Electrical Engineering and Computer Science	Publication Date: Jan 1, 2022
License type: cc-by-nc

Similar Papers

A novel density peaks clustering algorithm for mixed data
Mingjing Du ... Yu Xue
Pattern Recognition Letters | VOL. 97
Mingjing Du, et. al.Mingjing Du ... Yu Xue
03 Jul 2017
Pattern Recognition Letters | VOL. 97

Application of clustering strategy for automatic segmentation of tissue regions in mass spectrometry imaging.
Guang Xu ... Shengfeng Gan
Rapid communications in mass spectrometry : RCM | VOL. 38
Guang Xu, et. al.Guang Xu ... Shengfeng Gan
23 Feb 2024
Rapid communications in mass spectrometry : RCM | VOL. 38

Rough K-prototypes Clustering Algorithm Based on OTC Similarity and Between-Cluster Information for Mixed Data
Fumin Ma ... Ting Gong
-
Fumin Ma, et. al.Fumin Ma ... Ting Gong
15 Aug 2022
15 Aug 2022

An Affinity Propagation Clustering Algorithm for Mixed Numeric and Categorical Datasets
Kang Zhang ... Xingsheng Gu
Mathematical Problems in Engineering | VOL. 2014
Kang Zhang, et. al.Kang Zhang ... Xingsheng Gu
01 Jan 2014
Mathematical Problems in Engineering | VOL. 2014

Editage

Paperpal

R Discovery

Mind the Graph

R Discovery Prime

R Discovery Prime

Reclust: an efficient clustering algorithm for mixed data based on reclustering and cluster validation

Abstract

Talk to us

Similar Papers

More From: Indonesian Journal of Electrical Engineering and Computer Science